Fei He
7 indexed papers
Publications per year
Top categories
Frequent co-authors
Research Timeline
The paper introduces WebAgentGuard, a novel reasoning-driven, multimodal guard model that effectively detects prompt injection attacks in vulnerable web agents without compromising their functionality.
The paper proposes WARD, a robust and efficient defense model that secures web agents against prompt injection attacks embedded in web content, achieving high recall and low false positives even against adaptive attacks.
PRO-CUA introduces a process-reward optimization framework that enables efficient, step-level reinforcement learning for training computer use agents by decoupling environment interaction from policy optimization.
AliMark proposes a novel watermarking framework that treats sentence-level watermarking as a bit sequence alignment problem, significantly enhancing robustness against structural text perturbations like sentence splitting and merging.
AliMark proposes a novel framework that enhances the robustness of sentence-level watermarking by reformulating the problem as a bit sequence encoding and alignment task, significantly improving resilience against structural text perturbations like sentence splitting and merging.
This paper investigates the vulnerability of LLM-based automatic grading systems to prompt injection (PI) attacks, demonstrating that current systems are highly susceptible to manipulation that can lead to unfairly high scores.
This paper extends a previous framework for estimating neurophysiological connectivity using noise-driven heat modelling on graphs, adding regularisation for improved robustness and relaxing noise assumptions.
Papers
Connectivity Estimation using Stochastic Graph Heat Modelling
This paper extends a previous framework for estimating neurophysiological connectivity using noise-driven heat modelling on graphs, adding regularisation for improved robustness and relaxing noise ass…