20 results for “software review workflows”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
This paper presents an empirical study on how developers respond to agentic code reviews using CodeRabbit, revealing mixed reception and opportunities for improvement.
This paper investigates the prevalence and impact of relation chains in software review workflows using data from Gerrit, finding that they increase merge times and review effort propagation.
This paper proposes conversational AI review assistants for code review, systems that engage in conversation with developers instead of just generating comments.
The paper introduces a new structural merge tool that ensures parsability and universality in comparison to existing tools, resulting in fewer incorrect merge results.
This paper synthesizes practitioner discourse at scale to build a causal model explaining the impact of AI on code review, recovering mechanisms behind observed trends.
This paper empirically evaluates LLM-generated reviews for academic papers, finding that while LLM reviews show some alignment with human ones, authors can effectively 'game' the system using iterativ…
Neil Fendley, Zhengyu Liu, Aonan Guan, Jiacheng Zhong +1 more
The paper introduces JAW, a novel framework that demonstrates how adversaries can hijack agentic workflows on automation platforms like GitHub Actions by manipulating inputs based on context-grounded…
Shahbaz Siddeeq, Mateen Abbasi, Jussi Rasku, Zheying Zhang +3 more
This paper compares the quality and coverage of epic-organized LLM-generated Gherkin acceptance criteria with requirement-aligned generation, using four requirements documents from the PURE dataset.
Jun Zhang, JianYing Qu, Hanwen Du, Zhongkai Sun +2 more
The paper introduces Code-QA-Bench, a novel framework that rigorously separates genuine code reasoning from mere documentation memorization in repository-level code understanding benchmarks.
Phoenix is a multi-agent system that uses seven safety controls and a test evaluation strategy to resolve GitHub issues, achieving 75% oracle-resolution with no regressions on a curated benchmark and…
The paper introduces SafeAudit, a meta-audit framework that systematically enumerates test cases and uses a quantitative metric to uncover significant residual unsafe behaviors in LLM agents that exis…
This paper presents a software ecosystem enabling consistent development environments for running workflows across diverse hardware platforms.
QASecClaw, a multi-agent LLM system, significantly improves the accuracy of Static Application Security Testing (SAST) by using specialized LLM agents to filter out false positives, achieving an F1 sc…
This paper introduces tap, a file-based collaboration protocol enabling LLM agents from different vendors to collaborate on a shared codebase without shared memory or identical runtimes.
Jisen Li, Bingxuan Li, Nanyi Jiang, Xuying Ning +9 more
PaperPilot is an interactive literature search agent that constructs an executable DAG of paper-search operators based on user queries and feedback, improving search results and reducing errors.
This paper conducts rapid reviews of peer-reviewed publications to assess the availability of evidence on software engineering practices in government agencies.