Michael Qizhe Shieh
2 indexed papers
Publications per year
Top categories
Frequent co-authors
Research Timeline
This paper conducts the first real-world safety evaluation of the personal AI agent OpenClaw, demonstrating that its broad system access creates inherent vulnerabilities that significantly increase the attack success rate regardless of the underlying large language model.
This paper introduces RSIBench-Data, a controlled benchmark for evaluating data-centric research capabilities of LLM agents.
Papers
RSIBench-Data: Benchmarking Data-Centric Research for Recursive Self-Improvement
Fanqing Meng, Lingxiao Du, Qiguang Chen, Ziqi Zhao +3 more
This paper introduces RSIBench-Data, a controlled benchmark for evaluating data-centric research capabilities of LLM agents.