ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “relevance judgments”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.IREmpiricalRecentJul 17, 2026

LLMs Encode Relevance as a Layer-Wise Cross-Lingual Signal

Pietro Bernardelle, Samaneh Mohtadi, Stefano Civelli, Joel Mackenzie +1 more

This paper studies the linear decodability of query-document relevance from residual-stream activations in instruction-tuned large language models (LLMs) and compares it with generated relevance judgm…

View →
cs.IREmpiricalRecentJul 1, 2026

As It Was: Aligning LLM Search Evaluation with Historical User Preferences

Ali Vardasbi, Gustavo Penha, Enrico Palumbo, Claudia Hauff +2 more

This paper introduces a behavior-grounded Large Language Model (LLM) judge for evaluating search engine result pages, improving alignment with user preferences by up to 15% in a multilingual dataset.

View →
cs.IRcs.CLEmpiricalRecentJun 11, 2026

ADORE: Iterative Query Expansion with Retrieval-Grounded Relevance Feedback

Amin Bigdeli, Negar Arabzadeh, Radin Hamidi Rad, Sajad Ebrahimi +2 more

The paper introduces ADORE, an iterative framework for query expansion using LLMs, which turns retrieval outcomes into feedback for the next expansion.

View →
cs.IRcs.AIRecentMay 29, 2026

SPECTRA: Synthetic IR Test Collections with Relevance Oracles and Controlled Distractor Diagnostics

Eric Liang

The paper introduces SPECTRA, a scalable framework for generating large, synthetic, and controllable information retrieval test collections, demonstrating its ability to expose system scaling and fail…

View →
cs.AIcs.IRRecentMay 28, 2026

Rethinking Literature Search Evaluation: Deep Research Helps, and Human Citation Lists Are Not a Ground Truth

Gaurav Sahu, Laurent Charlin, Christopher Pal

The paper introduces a Deep Research pipeline that significantly improves literature search recall and demonstrates that human-curated citation lists are often unreliable and do not serve as a true gr…

View →
cs.IRcs.AIRecentMay 30, 2026

SkillPager: Query-Adaptive Intra-Skill Navigation via Semantic Node Retrieval

Zicai Cui, Zihan Guo, Weiwen Liu, Weinan Zhang

SkillPager is a novel two-stage framework that efficiently selects minimal, execution-sufficient context from large procedural skill documents by leveraging typed semantic nodes, significantly reducin…

View →
cs.AIRecentMay 27, 2026

Relevant Is Not Warranted: Evidence-Force Calibration for Cited RAG

Pin Qian, Su Wang, Xiaoyuan Wang, Yihang Chen +6 more

The paper introduces FORCEBENCH, a new stress test designed to evaluate whether cited sources genuinely warrant the strength of a claim, revealing that standard citation evaluation methods often fail…

View →
cs.IREmpiricalRecentJul 6, 2026

Prompting Beats Fine-Tuning: Generative Expected Value Scoring for Statutory Term Retrieval

Alvin Wang, Jaromir Savelka

The paper compares two families of methods for ranking case-law sentences by their usefulness for explaining statutory concepts using ModernBERT and decoder-only models. Decoder-only models achieve th…

View →
cs.CLcs.IRRecentJun 2, 2026

Re-Ranking Through an Attribution Lens for Citation Quality in Legal QA

Mohamed Hesham Elganayni, Selim Saleh

The paper introduces a cross-encoder re-ranker trained on attribution scores to improve the retrieval of highly relevant citation passages for legal question answering, outperforming standard semantic…

View →
cs.AIcs.IRRecentMay 28, 2026

Xetrieval: Mechanistically Explaining Dense Retrieval

Zhixin Cai, Jun Bai, Yang Liu, Jiaqi Li +6 more

Xetrieval introduces an embedding-level framework to mechanistically explain dense retrieval decisions by decomposing high-dimensional embeddings into sparse, human-interpretable features.

View →
cs.IREmpiricalRecentJul 21, 2026

An Epistemic Position-Based Click Model: From Interactions to Epistemic Distributions of Relevance and Bias

Oscar Rolando Ramirez Milian, Harrie Oosterhuis

This paper introduces the first evidential deep-learning approach for click modeling, providing beta-distributions for relevance and position-bias variables.

View →
cs.AIRecentMay 28, 2026

RAISE: RAG Design as an Architecture Search Problem

Zhen Chen, Yibing Liu, Weihao Xie, Yu Liang +2 more

The paper proposes formulating RAG design as an architecture search problem and introduces RAISE, a comprehensive framework and benchmark for systematically optimizing RAG hyperparameters.

View →
cs.CLcs.IREmpiricalRecentJun 10, 2026

uva-irlab-conv at SemEval-2026 Task 8: Multi-Turn RAG with Learned Sparse Retrieval and Listwise Reranking

Simon Lupart, Kidist Amde Mekonnen, Zahra Abbasiantaeb, Mohammad Aliannejadi

This paper proposes a multi-turn retrieval-augmented generation pipeline for conversational systems across four domains.

View →
cs.CLRecentMay 28, 2026

CanLegalRAGBench: Evaluating Retrieval-Augmented Generation on Canadian Case Law

Ethan Zhao, Maksym Taranukhin, Wei Cui, Moira Aikenhead +1 more

The paper introduces CanLegalRAGBench, a new Canadian legal QA benchmark, and evaluates RAG systems, finding that while open-source models are competitive, automatic evaluations struggle with nuanced…

View →
cs.CLRecentMay 31, 2026

Beyond Topical Similarity: Contrastive Evidence Retrieval with Interpretable Attention Alignment in RAG

Francielle Vargas, João Robiatti, Diego Alves, Lucas Pascotti Valem +5 more

The paper introduces CERA, a novel contrastive retrieval framework that improves RAG factuality and interpretability by using subjectivity-based hard negative selection and an auxiliary attention alig…

View →
cs.IREmpiricalRecentJul 7, 2026

Retrieving a Set, Not Independent Passages: Set-Level Compatibility Learning for Efficient Set Exploration

Mooho Song, Jay-Yoon Lee

This paper proposes a framework for multi-hop retrieval as query-set compatibility scoring, improving retrieval performance and downstream QA task performance.

View →
cs.IRcs.AIcs.CYRecentMay 27, 2026

Whose Name Comes Up? III: Persona Prompting Effects in LLM-Based Scholar Recommendation

Annabella Sánchez-Guzmán, Lukas Eberhard, Denis Helic, Lisette Espín-Noboa

The paper proposes a comprehensive benchmark to systematically audit how varying persona prompts and model choices affect the technical quality and social representativeness of scholar recommendations…

View →
cs.CLcs.AIRecentMay 27, 2026

UA-Legal-Bench: A Benchmark for Evaluating Large Language Models on Ukrainian Legal Reasoning

Volodymyr Ovcharov

The paper introduces UA-Legal-Bench, a comprehensive Ukrainian legal reasoning benchmark built from a massive judicial corpus, demonstrating that LLM performance is highly task-dependent and that simp…

View →
cs.IRcs.AIcs.CLEmpiricalRecentJul 16, 2026

Does generative AI supersede supervised XMLC? A Benchmark Study on Automated Subject Indexing with German Scientific Literature

Maximilian Kähler, Katja Konermann, Lisa Kluge, Markus Schumacher

This paper compares specialized supervised Extreme Multi-Label Classification (XMLC) methods with lexical matching baselines and LLM-based methods for subject indexing contemporary German scientific l…

View →