20 results for “relevance judgments”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
This paper studies the linear decodability of query-document relevance from residual-stream activations in instruction-tuned large language models (LLMs) and compares it with generated relevance judgm…
Ali Vardasbi, Gustavo Penha, Enrico Palumbo, Claudia Hauff +2 more
This paper introduces a behavior-grounded Large Language Model (LLM) judge for evaluating search engine result pages, improving alignment with user preferences by up to 15% in a multilingual dataset.
The paper introduces ADORE, an iterative framework for query expansion using LLMs, which turns retrieval outcomes into feedback for the next expansion.
The paper introduces SPECTRA, a scalable framework for generating large, synthetic, and controllable information retrieval test collections, demonstrating its ability to expose system scaling and fail…
The paper introduces a Deep Research pipeline that significantly improves literature search recall and demonstrates that human-curated citation lists are often unreliable and do not serve as a true gr…
SkillPager is a novel two-stage framework that efficiently selects minimal, execution-sufficient context from large procedural skill documents by leveraging typed semantic nodes, significantly reducin…
Pin Qian, Su Wang, Xiaoyuan Wang, Yihang Chen +6 more
The paper introduces FORCEBENCH, a new stress test designed to evaluate whether cited sources genuinely warrant the strength of a claim, revealing that standard citation evaluation methods often fail…
The paper compares two families of methods for ranking case-law sentences by their usefulness for explaining statutory concepts using ModernBERT and decoder-only models. Decoder-only models achieve th…
The paper introduces a cross-encoder re-ranker trained on attribution scores to improve the retrieval of highly relevant citation passages for legal question answering, outperforming standard semantic…
Zhixin Cai, Jun Bai, Yang Liu, Jiaqi Li +6 more
Xetrieval introduces an embedding-level framework to mechanistically explain dense retrieval decisions by decomposing high-dimensional embeddings into sparse, human-interpretable features.
This paper introduces the first evidential deep-learning approach for click modeling, providing beta-distributions for relevance and position-bias variables.
Zhen Chen, Yibing Liu, Weihao Xie, Yu Liang +2 more
The paper proposes formulating RAG design as an architecture search problem and introduces RAISE, a comprehensive framework and benchmark for systematically optimizing RAG hyperparameters.
This paper proposes a multi-turn retrieval-augmented generation pipeline for conversational systems across four domains.
Ethan Zhao, Maksym Taranukhin, Wei Cui, Moira Aikenhead +1 more
The paper introduces CanLegalRAGBench, a new Canadian legal QA benchmark, and evaluates RAG systems, finding that while open-source models are competitive, automatic evaluations struggle with nuanced…
The paper introduces CERA, a novel contrastive retrieval framework that improves RAG factuality and interpretability by using subjectivity-based hard negative selection and an auxiliary attention alig…
This paper proposes a framework for multi-hop retrieval as query-set compatibility scoring, improving retrieval performance and downstream QA task performance.
The paper proposes a comprehensive benchmark to systematically audit how varying persona prompts and model choices affect the technical quality and social representativeness of scholar recommendations…
The paper introduces UA-Legal-Bench, a comprehensive Ukrainian legal reasoning benchmark built from a massive judicial corpus, demonstrating that LLM performance is highly task-dependent and that simp…
This paper compares specialized supervised Extreme Multi-Label Classification (XMLC) methods with lexical matching baselines and LLM-based methods for subject indexing contemporary German scientific l…