Hao Peng
6 indexed papers
Publications per year
Top categories
Frequent co-authors
Research Timeline
SafeHarbor is a novel, hierarchical memory-augmented framework that establishes context-aware decision boundaries for LLM agents, achieving state-of-the-art safety while minimizing over-refusal.
The paper proposes using an auxiliary reconstruction task, specifically one that captures intra-state feature dependencies, to improve the quality of state representations learned by the encoder in neural algorithmic reasoning.
This paper introduces CHERRL, a controllable hacking environment for rubric-based reinforcement learning to study and mitigate reward hacking.
The paper proposes OneReason, a framework that enhances the reasoning capability of generative recommendation models by focusing on improving item perception and structuring user behavior into coherent latent interests.
This paper proposes a method for creating scalable browser agents by cloning user interaction skills from human browsing data using natural language skills and a skill graph.
A new framework, PRTA, is proposed for full-ranking recommendation tasks using large language models, where an LLM acts as a central planner and traditional recommendation models perform scoring.
Papers
Personalized Recommendation Tool Learning via Autonomous Language Agents
Mingdai Yang, Zhiwei Liu, Weizhi Zhang, Yibo Wang +2 more
A new framework, PRTA, is proposed for full-ranking recommendation tasks using large language models, where an LLM acts as a central planner and traditional recommendation models perform scoring.