Yue Huang
3 indexed papers
Publications per year
Top categories
Frequent co-authors
Research Timeline
The paper introduces AgentTrap, a dynamic benchmark that measures LLM agent susceptibility to malicious side effects embedded within seemingly benign third-party skills, finding that agents often execute unsafe side effects while completing the visible user task.
The paper introduces a Contextual Integrity (CI) framework and a new benchmark (DelegateCI-Bench) to rewrite user queries sent to cloud LLMs, ensuring only task-essential information is retained while preserving utility and maximizing privacy.
This paper investigates safety dynamics in the use of Role-play AI Companions through interviews and a 14-day assessment, identifying key factors shaping these dynamics and revealing short-term emotional relief masking longer-term deterioration.
Papers
Beyond Her: Safety Dynamics in Role-play AI Companions
Zehang Deng, Zhaoyang Xie, Changzhou Han, Hiran Thabrew +7 more
This paper investigates safety dynamics in the use of Role-play AI Companions through interviews and a 14-day assessment, identifying key factors shaping these dynamics and revealing short-term emotio…