ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “review effort”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.SEEmpiricalRecentJul 22, 2026

How Developers Use Relation Chains in Gerrit-Based Review Ecosystems: An Empirical Study Across Three Open-Source Ecosystems

Ahmed Belhouchette, Moataz Chouchen, Marouene Chaieb, Mohammad Hamdaqa Abdelwahab Hamou-Lhadj

This paper investigates the prevalence and impact of relation chains in software review workflows using data from Gerrit, finding that they increase merge times and review effort propagation.

View →
cs.AIRecentMay 27, 2026

Relevant Is Not Warranted: Evidence-Force Calibration for Cited RAG

Pin Qian, Su Wang, Xiaoyuan Wang, Yihang Chen +6 more

The paper introduces FORCEBENCH, a new stress test designed to evaluate whether cited sources genuinely warrant the strength of a claim, revealing that standard citation evaluation methods often fail…

View →
cs.AIcs.MARecentMay 27, 2026

Review Arcade: On the Human Alignment and Gameability of LLM Reviews

Hans Ole Hatzel, Sebastian Steindl, Jan Strich

This paper empirically evaluates LLM-generated reviews for academic papers, finding that while LLM reviews show some alignment with human ones, authors can effectively 'game' the system using iterativ…

View →
cs.LGcs.AIcs.CLEmpiricalRecentJun 26, 2026

Towards Automating Scientific Review with Google's Paper Assistant Tool

Rajesh Jayaram, Drew Tyler, David Woodruff, Corinna Cortes +3 more

The paper proposes a taxonomy for AI-human collaboration in scientific evaluation and introduces the Paper Assistant Tool (PAT) to accelerate verification and review process in scientific research.

View →
cs.AIcs.ARRecentMay 31, 2026

Can AI Review Improve Paper Drafting? An Empirical Study on 20 Computer Architecture Submissions

Di Wu

The paper empirically investigates whether AI-generated reviews can improve the drafting process of academic papers, finding that AI reviews cover many human-identified issues but also introduce novel…

View →
cs.SEcs.AIEmpiricalRecentJul 8, 2026

3100 Opinions on Code Review in an AI World: Building Causal Theory from Practitioner Discourse

Shyam Agarwal, Courtney Miller, Christian Kästner, Bogdan Vasilescu

This paper synthesizes practitioner discourse at scale to build a causal model explaining the impact of AI on code review, recovering mechanisms behind observed trends.

View →
cs.AIcs.CLcs.CYRecentMay 29, 2026

On Wednesdays, We Ask Questions: Optimizing "Active Listening" in Automated Legal Triage and Referral

Quinten Steenhuis, Jacqueline Harvey

The paper evaluates an automated legal triage system (FETCH) that uses follow-up questions, demonstrating that while low-cost LLMs are effective for classification, generating high-quality questions r…

View →
cs.SEcs.AIEmpiricalRecentJul 3, 2026

Is Agentic Code Review Helpful? Mining Developers' Feedback to CodeRabbit Reviews in the Wild

Hong Yi Lin, Mingzhao Liang, Kla Tantithamthavorn, Patanamon Thongtanunam

This paper presents an empirical study on how developers respond to agentic code reviews using CodeRabbit, revealing mixed reception and opportunities for improvement.

View →
cs.SEPositionRecentJul 24, 2026

Code Review is a Conversation: Toward Conversational AI Review Assistants

Rosalia Tufano

This paper proposes conversational AI review assistants for code review, systems that engage in conversation with developers instead of just generating comments.

View →
cs.LGcs.AIRecentMay 27, 2026

FormInv: A Measurement Protocol for Semantic Invariance in Mathematical Reasoning Benchmarks

Nishal Thomas, Noel Thomas

The paper introduces FormInv, a measurement protocol that reveals significant semantic inconsistencies in existing mathematical reasoning benchmarks, showing that standard accuracy metrics fail to cap…

View →
cs.AIcs.CLRecentMay 28, 2026

PRAIB: Peer Review AI Benchmark of Behaviour of LLM-Assisted Reviewing

Krzysztof Żurawicki, Julia Farganus, Arkadiusz Gaweł, Mateusz Bystroński +1 more

The paper introduces PRAIB, a benchmark that demonstrates that LLM-generated peer reviews, while often verbose, systematically diverge from human norms by being less variable, positively biased, and f…

View →
cs.CLcs.IREmpiricalRecentJul 15, 2026

DS@GT ARC at LongEval: Citation Integrity and Factual Grounding in Scientific QA

Brandon Michaels, Brendon Johnson

DS@GT ARC evaluates the effectiveness of Corrective RAG (CRAG) and CiteFix in improving citation faithfulness and answer grounding in Retrieval-Augmented Generation (RAG) QA systems, revealing a trade…

View →
cs.CLRecentMay 28, 2026

Auditing LLM Benchmarks with Item Response Theory

Sander Land, Daniel M. Bikel

The paper introduces an Item Response Theory (IRT)-based indicator that effectively identifies likely mislabeled items in existing LLM benchmarks, revealing systematic errors in labeling and model spe…

View →
cs.DCcs.AIcs.LGRecentMay 31, 2026

Hierarchical Online Prompt Mutation with Dual-Loop Feedback for Guardrailed Evidence Document Generation: A Production-Evaluation Case Study

Nataraj Agaram Sundar Tejas Morabia

The paper introduces HOPM, a hierarchical online prompt mutation framework that significantly improves the performance of language models in high-stakes evidence document generation by integrating dua…

View →
cs.AIRecentMay 27, 2026

Better Accuracies, Worse Reasoning: A Step-Level Audit of Medical Chain-of-Thought Distillation

Zhaoyang Jiang, Xuanqi Peng, Fei Teng, Zhizhong Fu +4 more

The paper demonstrates that while distilling large language models for medical QA can significantly improve final answer accuracy, this gain often comes at the cost of factual accuracy and detailed re…

View →
cs.CLcs.AIcs.IRRecentMay 27, 2026

Same Question, Different Source, Different Answer: Auditing Source-Dependence in Medical Multi-Source RAG

Yubo Li, Rema Padman, Ramayya Krishnan

This paper introduces a framework to audit source-dependence in multi-source RAG systems, demonstrating that disagreement across institutional sources is a common and critical failure mode that curren…

View →
cs.CLcs.AIEmpiricalRecentJun 26, 2026

Can LLMs Judge Better Than They Generate? Evaluating Task Asymmetry, Mechanistic Interpretability and Transferability for In-Context QA

Sambaran Bandyopadhyay

This paper tests the assumption that evaluation is easier than generation in LLM-as-a-Judge and self-evaluation pipelines using a controlled in-context QA setting and reveals that evaluation attends t…

View →
cs.CLcs.AIRecentMay 28, 2026

CRITIC-R1: Learning Structured Critics for Retrieval-Augmented Generation

Wenhan Xiao, Ziwei Zhang, Chuanyue Yu, Xingcheng Fu +3 more

CRITIC-R1 introduces a structured critic framework that treats RAG critique as an explicit error diagnosis problem using reinforcement learning, significantly improving answer quality over strong RAG…

View →