ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “software review workflows”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.SEcs.AIEmpiricalRecentJul 3, 2026

Is Agentic Code Review Helpful? Mining Developers' Feedback to CodeRabbit Reviews in the Wild

Hong Yi Lin, Mingzhao Liang, Kla Tantithamthavorn, Patanamon Thongtanunam

This paper presents an empirical study on how developers respond to agentic code reviews using CodeRabbit, revealing mixed reception and opportunities for improvement.

View →
cs.SEEmpiricalRecentJul 22, 2026

How Developers Use Relation Chains in Gerrit-Based Review Ecosystems: An Empirical Study Across Three Open-Source Ecosystems

Ahmed Belhouchette, Moataz Chouchen, Marouene Chaieb, Mohammad Hamdaqa Abdelwahab Hamou-Lhadj

This paper investigates the prevalence and impact of relation chains in software review workflows using data from Gerrit, finding that they increase merge times and review effort propagation.

View →
cs.SEPositionRecentJul 24, 2026

Code Review is a Conversation: Toward Conversational AI Review Assistants

Rosalia Tufano

This paper proposes conversational AI review assistants for code review, systems that engage in conversation with developers instead of just generating comments.

View →
cs.SEEmpiricalRecentJul 8, 2026

On the Correctness of Software Merge

Akira Mori, Masatomo Hashimoto

The paper introduces a new structural merge tool that ensures parsability and universality in comparison to existing tools, resulting in fewer incorrect merge results.

View →
cs.SEcs.AIEmpiricalRecentJul 8, 2026

3100 Opinions on Code Review in an AI World: Building Causal Theory from Practitioner Discourse

Shyam Agarwal, Courtney Miller, Christian Kästner, Bogdan Vasilescu

This paper synthesizes practitioner discourse at scale to build a causal model explaining the impact of AI on code review, recovering mechanisms behind observed trends.

View →
cs.AIcs.MARecentMay 27, 2026

Review Arcade: On the Human Alignment and Gameability of LLM Reviews

Hans Ole Hatzel, Sebastian Steindl, Jan Strich

This paper empirically evaluates LLM-generated reviews for academic papers, finding that while LLM reviews show some alignment with human ones, authors can effectively 'game' the system using iterativ…

View →
cs.CRcs.AIcs.SERecentMay 11, 2026

Comment and Control: Hijacking Agentic Workflows via Context-Grounded Evolution

Neil Fendley, Zhengyu Liu, Aonan Guan, Jiacheng Zhong +1 more

The paper introduces JAW, a novel framework that demonstrates how adversaries can hijack agentic workflows on automation platforms like GitHub Actions by manipulating inputs based on context-grounded…

View →
cs.SEEmpiricalRecentJul 2, 2026

Epic-Organized vs. Requirement-Aligned Gherkin: An Empirical Evaluation of LLM-Based Acceptance Criteria Generation

Shahbaz Siddeeq, Mateen Abbasi, Jussi Rasku, Zheying Zhang +3 more

This paper compares the quality and coverage of epic-organized LLM-generated Gherkin acceptance criteria with requirement-aligned generation, using four requirements documents from the PURE dataset.

View →
cs.SEcs.AIRecentMay 28, 2026

Code-QA-Bench: Separating Code Reasoning from Documentation Memorization in Repository-Level QA

Jun Zhang, JianYing Qu, Hanwen Du, Zhongkai Sun +2 more

The paper introduces Code-QA-Bench, a novel framework that rigorously separates genuine code reasoning from mere documentation memorization in repository-level code understanding benchmarks.

View →
cs.SEcs.MAEmpiricalRecentJun 18, 2026

Phoenix: Safe GitHub Issue Resolution via Multi-Agent LLMs

Kipngeno Koech, Muhammad Adam, Baimam Boukar Jean Jacques, Joao Barros

Phoenix is a multi-agent system that uses seven safety controls and a test evaluation strategy to resolve GitHub issues, achieving 75% oracle-resolution with no regressions on a curated benchmark and…

View →
cs.SEcs.CRRecentMar 18, 2026

Who Tests the Testers? Systematic Enumeration and Coverage Audit of LLM Agent Tool Call Safety

Xuan Chen, Lu Yan, Ruqi Zhang, Xiangyu Zhang

The paper introduces SafeAudit, a meta-audit framework that systematically enumerates test cases and uses a quantitative metric to uncover significant residual unsafe behaviors in LLM agents that exis…

View →
cs.DCEmpiricalRecentJul 21, 2026

A User-oriented Portable, Reproducible, and Scalable Software Ecosystem

Alfio Lazzaro, Utz-Uwe Haus, Sandrine Charousset, Nina Mujkanovic

This paper presents a software ecosystem enabling consistent development environments for running workflows across diverse hardware platforms.

View →
cs.CRcs.SERecentMay 3, 2026

QASecClaw: A Multi-Agent LLM Approach for False Positive Reduction in Static Application Security Testing

Mohd Ruhul Ameen, Md Takrim Ul Alam, Akif Islam

QASecClaw, a multi-agent LLM system, significantly improves the accuracy of Static Application Security Testing (SAST) by using specialized LLM agents to filter out false positives, achieving an F1 sc…

View →
cs.SEcs.AIcs.HCEmpiricalRecentJun 12, 2026

tap: A File-Based Protocol for Heterogeneous LLM Agent Collaboration

Minseo Kim

This paper introduces tap, a file-based collaboration protocol enabling LLM agents from different vendors to collaborate on a shared codebase without shared memory or identical runtimes.

View →
cs.CLcs.IREmpiricalRecentJul 1, 2026

Multi-Turn Agentic Scientific Literature Search via Workflow Induction

Jisen Li, Bingxuan Li, Nanyi Jiang, Xuying Ning +9 more

PaperPilot is an interactive literature search agent that constructs an executable DAG of paper-search operators based on user queries and feedback, improving search results and reducing errors.

View →
cs.SEcs.DLEmpiricalRecentJul 24, 2026

A Preliminary Search for Evidence on Government Software Engineering Practices: Results from Three Rapid Reviews

Sebastián Pizard, Matías Porro, Andrea Muñoz, Andrea Delgado

This paper conducts rapid reviews of peer-reviewed publications to assess the availability of evidence on software engineering practices in government agencies.

View →