ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “academic search”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.IRcs.HCEmpiricalRecentJul 3, 2026

AI Overviews in Academic Search: Evaluating AI-generated Summaries of Search Results in a Domain-specific Search Engine

Kevin Schott, Kanishka Silva, Ingo Frommholz, Philipp Mayr +2 more

This paper evaluates the use of AI-generated summaries on search engine results pages (SERPs) in academic search for social science information.

View →
cs.IRcs.AIcs.CLEmpiricalRecentJul 2, 2026

Evaluating Chunking Strategies for Retrieval-Augmented Generation on Academic Texts

Valentin J. J. Kreileder, Johannes Reisinger, Andreas Fischer

This paper evaluates the effectiveness of cluster-based semantic chunking compared to fixed-size and recursive chunking in Retrieval-Augmented Generation systems using the Retrieval Augmented Generati…

View →
cs.AIRecentMay 27, 2026

LiveBrowseComp: Are Search Agents Searching, or Just Verifying What They Already Know?

HuiMing Fan, Xiao Wang, Zheng Chu, Qianyu Wang +4 more

The paper argues that current search agents often verify existing knowledge rather than genuinely searching, and introduces LiveBrowseComp, a new benchmark to measure true evidence-driven discovery.

View →
cs.AIRecentMay 28, 2026

RAISE: RAG Design as an Architecture Search Problem

Zhen Chen, Yibing Liu, Weihao Xie, Yu Liang +2 more

The paper proposes formulating RAG design as an architecture search problem and introduces RAISE, a comprehensive framework and benchmark for systematically optimizing RAG hyperparameters.

View →
cs.IRRecentJun 3, 2026

SearchLog: A Web Browser Extension for Capturing Search Logs in Laboratory Studies

Jiaman He, Riccardo Xia, Dana McKay, Damiano Spina +1 more

The paper presents SearchLog, a web browser extension for collecting natural search logs during lab-based studies.

View →
cs.AIcs.IRRecentMay 28, 2026

Rethinking Literature Search Evaluation: Deep Research Helps, and Human Citation Lists Are Not a Ground Truth

Gaurav Sahu, Laurent Charlin, Christopher Pal

The paper introduces a Deep Research pipeline that significantly improves literature search recall and demonstrates that human-curated citation lists are often unreliable and do not serve as a true gr…

View →
cs.CLcs.IREmpiricalRecentJul 1, 2026

Multi-Turn Agentic Scientific Literature Search via Workflow Induction

Jisen Li, Bingxuan Li, Nanyi Jiang, Xuying Ning +9 more

PaperPilot is an interactive literature search agent that constructs an executable DAG of paper-search operators based on user queries and feedback, improving search results and reducing errors.

View →
cs.IRcs.AIcs.CYRecentMay 27, 2026

Whose Name Comes Up? III: Persona Prompting Effects in LLM-Based Scholar Recommendation

Annabella Sánchez-Guzmán, Lukas Eberhard, Denis Helic, Lisette Espín-Noboa

The paper proposes a comprehensive benchmark to systematically audit how varying persona prompts and model choices affect the technical quality and social representativeness of scholar recommendations…

View →
cs.IREmpiricalRecentJun 11, 2026

Hybrid Neural Retrieval with Generative Query Refinement for Quranic Passage Retrieval

Mohamed G. Salman, Mohammad E. Moftah, Ali Hamdi

This paper proposes a neural architecture for Quranic Passage Retrieval using hybrid candidate retrieval, semantic reranking, confidence gating, and multi-verse aggregation.

View →
cs.IRcs.AIRecentMay 29, 2026

SPECTRA: Synthetic IR Test Collections with Relevance Oracles and Controlled Distractor Diagnostics

Eric Liang

The paper introduces SPECTRA, a scalable framework for generating large, synthetic, and controllable information retrieval test collections, demonstrating its ability to expose system scaling and fail…

View →
cs.AIcs.CLRecentMay 27, 2026

A Fixed-Budget, Cluster-Aware Standard for LLM-as-a-Judge Evaluation: A Multi-Hop RAG Stress Test

Camilo Chacón Sartori, José H. García

The paper proposes a rigorous, fixed-budget, cluster-aware standard for LLM-as-a-judge evaluation of multi-hop RAG systems, demonstrating that current evaluation methods often overstate performance.

View →
cs.IREmpiricalRecentJul 1, 2026

As It Was: Aligning LLM Search Evaluation with Historical User Preferences

Ali Vardasbi, Gustavo Penha, Enrico Palumbo, Claudia Hauff +2 more

This paper introduces a behavior-grounded Large Language Model (LLM) judge for evaluating search engine result pages, improving alignment with user preferences by up to 15% in a multilingual dataset.

View →
cs.CLcs.AIcs.IRRecentMay 28, 2026

GrepSeek: Training Search Agents for Direct Corpus Interaction

Alireza Salemi, Chang Zeng, Atharva Nijasure, Jui-Hui Chung +3 more

GrepSeek introduces a novel direct corpus interaction (DCI) search agent that trains an LLM to find and compose evidence from large text corpora by issuing executable shell commands, achieving state-o…

View →
cs.IRcs.AIcs.LGRecentMay 28, 2026

No More K-means: Single-Stage Sparse Coding for Efficient Multi-Vector Retrieval

Lixuan Guo, Yifei Wang, Tiansheng Wen, Aosong Feng +2 more

The paper introduces Single-stage Sparse Retrieval (SSR), a method that replaces computationally expensive vector clustering with sparse autoencoding to achieve highly efficient multi-vector retrieval…

View →
cs.IREmpiricalRecentJul 17, 2026

Scientific Claim-Source Retrieval Revisited: A Comparative Study of Style Transfer and Re-Ranking

Tobias Schreieder, Harsh Khandelwal, Yu-Ling Zhong, Michael Färber

This paper compares sparse and dense retrieval models for scientific claim-source retrieval on the CheckThat! 2026 benchmark. Translating claims into English and incorporating publication metadata imp…

View →
cs.IRcs.CLEmpiricalRecentJul 9, 2026

Improving Ad-hoc Search Effectiveness for Conversational Information Retrieval via Model Merging

Ahmed Rayane Kebir, Jose G. Moreno, Lynda Tamine

This paper introduces model merging as a training-free strategy for designing a single retrieval model that operates across both ad-hoc and conversational settings, improving ad-hoc search capabilitie…

View →
cs.AIcs.IRRecentMay 27, 2026

From Learning Resources to Competencies: LLM-Based Tagging with Evidence and Graph Constraints

Ngoc Luyen Le, Marie-Hélène Abel, Bertrand Laforge

The paper introduces an LLM-based pipeline that tags learning resources with structured competencies, achieving strong performance while providing traceable evidence and leveraging graph constraints.

View →
cs.IREmpiricalRecentJul 21, 2026

TSGR: Taobao Search Generative Retrieval

Tianyu Zhan, Gui Ling, Tong Xiong, Kunhai Lin +8 more

This paper proposes TSGR, a generative retrieval framework for industrial e-commerce search that incorporates value awareness into item representation and candidate ranking.

View →
cs.IREmpiricalRecentJul 22, 2026

CIR at iKAT SCAI 2026: Exploring Clarification Need Prediction in Agentic Conversational Search

Nolwenn Bernard, Jüri Keller, Philipp Schaer

The Cologne Information Retrieval group participated in iKAT SCAI 2026 shared task using an agentic conversational search system with query rewriting, retrieval, reranking, answer generation, and clar…

View →