ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “Familiarity with question answering tasks”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

stat.OTcs.AIEmpiricalRecentJun 9, 2026

Flaws in the LLM Automation Narrative

George Perrett, Javae Elliott, Jennifer Hill, Marc Scott

This paper evaluates the performance of a Large Language Model (LLM) in a high-stakes context by comparing it to human experts and measuring variance and error magnitude.

View →
cs.CLcs.IREmpiricalRecentJun 10, 2026

uva-irlab-conv at SemEval-2026 Task 8: Multi-Turn RAG with Learned Sparse Retrieval and Listwise Reranking

Simon Lupart, Kidist Amde Mekonnen, Zahra Abbasiantaeb, Mohammad Aliannejadi

This paper proposes a multi-turn retrieval-augmented generation pipeline for conversational systems across four domains.

View →
cs.IREmpiricalRecentJul 7, 2026

Retrieving a Set, Not Independent Passages: Set-Level Compatibility Learning for Efficient Set Exploration

Mooho Song, Jay-Yoon Lee

This paper proposes a framework for multi-hop retrieval as query-set compatibility scoring, improving retrieval performance and downstream QA task performance.

View →
cs.CLRecentJun 1, 2026

Not What, But How: A Communicative Audit of LLM Response Framing

Siddhesh Milind Pawar, Sarah Masud, Haneul Yoo, Alice Oh +1 more

The paper introduces FRANZ, a communicative audit framework, to evaluate how LLMs frame responses to subjective questions, finding that LLMs exhibit statistically significant and coupled differences i…

View →
cs.CLEmpiricalRecentJul 23, 2026

When Trivia Is Not Trivial: Everyday Knowledge Failures in Multilingual LLMs

Anna Mosolova, Djamé Seddah

This paper evaluates the performance of large language models on quiz-style questions covering common and niche topics in six European languages, finding significant knowledge gaps and language-depend…

View →
cs.CLEmpiricalRecentJul 2, 2026

CheckRLM: Effective Knowledge-Thought Coherence Checking in Retrieval-Augmented Reasoning

Dingling Xu, Ruobing Wang, Qingfei Zhao, Yukun Yan +7 more

The paper proposes CheckRLM, a framework that improves the reliability of Reasoning Language Models by identifying and correcting factual errors using Retrieval-Augmented Generation.

View →
cs.CLRecentMay 30, 2026

OCC-RAG: Optimal Cognitive Core for Faithful Question Answering

Maksim Savkin, Mikhail Goncharov, Alexander Gambashidze, Alla Chepurova +6 more

The paper introduces OCC-RAG, a family of compact, task-specialized Small Language Models (SLMs) designed to achieve highly faithful, multi-hop question answering grounded strictly in provided context…

View →
cs.CLcs.IREmpiricalRecentJul 15, 2026

DS@GT ARC at LongEval: Citation Integrity and Factual Grounding in Scientific QA

Brandon Michaels, Brendon Johnson

DS@GT ARC evaluates the effectiveness of Corrective RAG (CRAG) and CiteFix in improving citation faithfulness and answer grounding in Retrieval-Augmented Generation (RAG) QA systems, revealing a trade…

View →
cs.IRcs.AIcs.CLRecentMay 29, 2026

On the impact of retrieved content representations in RAG Pipelines

Jonathan J Ross, Bevan Koopman, Anton van der Vegt, Guido Zuccon

The paper systematically compares multiple content representations for RAG pipelines and finds that answer retention—the ability of the representation to preserve the original answer-bearing content—i…

View →
cs.CLcs.AIcs.IREmpiricalRecentJun 27, 2026

AB-RAG: Adaptive Budgeted Retrieval-Augmented Generation for Reliable Question Answering

Ansh Kamthan

This paper introduces AB-RAG, a training-free and backbone-agnostic framework for adaptively generating answers, estimating their confidence, and deciding whether to retrieve more evidence based on th…

View →
cs.CLRecentMay 29, 2026

What Am I Missing? Question-Answering as Hidden State Probing

Chu Fei Luo, Samuel Dahan, Xiaodan Zhu

The paper proposes using question-asking as an inference-time intervention to probe a language model's hidden state, finding that the self-diagnosis process provides a predictive signal for final corr…

View →
cs.IREmpiricalRecentJul 24, 2026

The Prompt Is Not the Query: How Request State Evolves Across Multi-Turn AI Conversations

Benjamin Tannenbaum

This paper investigates how the final prompt in conversational AI-search evaluations differs from the conversation history, using two corpora of commercial and PRISM conversations.

View →
cs.CLcs.AIRecentMay 31, 2026

Connecting the Dots: Benchmarking Reflective Memory in Long-Horizon Dialogue

Jingjie Lin, Bingbing Wang, Zihan Wang, Zhengda Jin +3 more

The paper introduces RefMem-Bench, a new benchmark for measuring reflective memory in long-horizon dialogue, and proposes REMIND, a framework that significantly improves models' ability to synthesize…

View →
cs.CLcs.AIRecentMay 30, 2026

SPADER: Step-wise Peer Advantage with Diversity-Aware Exploration Rewards for Multi-Answer Question Answering

Qiming Shi, Zhaolu Kang, Yunfan Zhou, Di Weng +1 more

SPADER is a novel reinforcement learning framework that addresses the challenges of Multi-Answer Question Answering by improving credit assignment and promoting diverse exploration during long-horizon…

View →
cs.CLcs.IRRecentMay 29, 2026

Beyond Static Dialogues: Benchmarking Realistic, Heterogeneous, and Evolving Long-Term Memory

Han Zhang, Zihao Tang, Xin Yu, Xiao Liu +7 more

The paper introduces RHELM, a new benchmark designed to test LLMs' long-term memory by simulating realistic, complex, and evolving dialogues that integrate multiple heterogeneous data sources.

View →
eess.ASEmpiricalRecentJul 21, 2026

Summary of DCASE 2026 Task 5: Audio-Dependent Question Answering

Haolin He, Renhe Sun, Zheqi Dai, Xingjian Du +15 more

This paper introduces Audio-Dependency Filtering (ADF) pipeline for Audio-Dependent Question Answering (ADQA) task in DCASE~2026, achieving top overall and sub-10B accuracy.

View →
cs.CLRecentMay 31, 2026

Not All Explanations Simulate Equally: Comparing Verbalized Feature Attributions and Self-Generated Rationales

Pingjun Hong, Benjamin Roth

The paper compares verbalized feature attributions and self-generated rationales for explaining model behavior, finding that the format and granularity of the explanation significantly affect its abil…

View →
cs.CLcs.AIRecentMay 31, 2026

CA-BED: Conversation-Aware Bayesian Experimental Design

Daniel Arnould, Rashad Aziz, Zixuan Kang, Tanav Changal +4 more

CA-BED is a novel framework that improves LLM performance in interactive question-answering by integrating Bayesian Experimental Design to strategically select questions that maximize information gain…

View →
cs.CLcs.AIEmpiricalRecentJul 23, 2026

RUMBA: Russian User Memory Benchmark

Elizaveta Shevtsova, Inna Glebkina, Mark Baushenko, Pavel Gulyaev +1 more

The paper introduces RUMBA, a new benchmark for long-term conversational memory in LLMs with a fine-grained taxonomy and unified methodology.

View →