Bo Li
30 indexed papers
Publications per year
Top categories
Frequent co-authors
Research Timeline
The paper introduces Kumushi, a root-cause-driven patching agent that significantly improves automated vulnerability repair by focusing LLMs on the true source of bugs, outperforming existing methods and matching commercial agents.
The paper introduces TurnGate, a response-aware defense mechanism that detects the earliest turn in a multi-turn dialogue where the accumulated interaction enables a harmful action, significantly improving malicious intent detection.
The paper proposes WARD, a robust and efficient defense model that secures web agents against prompt injection attacks embedded in web content, achieving high recall and low false positives even against adaptive attacks.
The paper identifies a failure mode called unfaithful capitulation (UC), where reasoning models maintain a correct internal thought process (chain-of-thought) but output an incorrect final answer when subjected to sustained adversarial questioning.
This paper introduces a framework to audit source-dependence in multi-source RAG systems, demonstrating that disagreement across institutional sources is a common and critical failure mode that current evaluation metrics overlook.
The paper proposes using GUI agents, both as objective evaluators and subjective playtesters, to significantly improve the generation of playable games from prompts, demonstrating a 66.8% rubric pass-rate with a novel iterative framework.
The paper introduces Loong, a novel human-like agent that significantly improves long document translation by adaptively selecting and utilizing optimal historical context using a specialized memory module and reinforcement learning.
The paper introduces a quotient-DAG view to accurately estimate unordered slate propensities for off-policy evaluation, solving the nuisance variance and computational gap inherent in standard importance sampling for autoregressive recommenders.
The paper introduces DrugClaw, a multi-agent system, and DrugAudit, a new benchmark, demonstrating that DrugClaw excels at answering drug-related questions by grounding answers in primary regulatory sources.
The paper introduces Diversity-inducing Initialization (DivIn), a novel method that improves image diversity by re-weighting the initial noise selection based on the guidance potential, thereby mitigating mode collapse.
This paper proposes a hybrid two-stage diffusion transformer architecture for instruction-guided audio editing, balancing performance and efficiency.
The paper introduces PATH, an in-situ indexing architecture for Processing-in-Memory systems that achieves higher throughput, lower tail latency, and fewer memory accesses than state-of-the-art schemes.
This paper accelerates conformal prediction by incorporating approximate leave-one-out estimators and establishes asymptotic coverage and efficiency.
This paper proposes SkillOpt-Lite, a minimal viable pipeline for skill optimization in autonomous agents, which accelerates convergence and outperforms full SkillOpt.
This paper introduces HoloGeo, an evidence-driven reasoning framework to mitigate landmark bias in Vision-Language Models, and establishes metrics and a benchmark to evaluate its effectiveness.
MagicSelector is a framework for tool retrieval in agents using counterfactual task decomposition, progressive reranking, and dynamic Top-K.
This paper proposes MineValiCoder, a collaborative closed-loop TDD framework using mutual reinforcement of test-case quality and code quality to address stochasticity in Large Language Model-based Test-Driven Development.
The paper introduces DBA-Bench, a benchmark for evaluating database agents with production fidelity, outcome-first evaluation, and controlled scenario reproducibility.
This paper introduces the Organizational Consensus Algorithm (OCA) for internal negotiation and decision coordination in organizational structures, modeling inter-departmental conflict as a dynamic game and using a retrospective penalty system.
This paper introduces AgentSysBench, a benchmark suite and measurement toolkit for agentic applications, and identifies six properties that distinguish agentic workloads from conventional LLM serving.
Papers
From LLM Inference to Agentic Workloads: Characterization and Implications for Serving Systems
Chaokun Chang, Yukun Zhou, Kaihua Fu, Dakai An +9 more
This paper introduces AgentSysBench, a benchmark suite and measurement toolkit for agentic applications, and identifies six properties that distinguish agentic workloads from conventional LLM serving.