Haiwen Hong
3 indexed papers
Publications per year
Top categories
Frequent co-authors
Research Timeline
The paper quantifies the exact parametric memory capacity of LLMs using LoRA and proposes a new optimization strategy, MemFT, to enhance memory fidelity.
This paper proposes Perceive-to-Reason (P2R), a framework for fine-grained visual reasoning that decouples perception from reasoning and introduces a new reinforcement learning strategy.
The paper introduces Relay On-Policy Distillation (Relay-OPD), a method for on-policy distillation that constructs relay trajectories to address prefix failure and improve performance.
Papers
Pass the Baton: Trajectory-Relayed On-Policy Distillation
Haolei Xu, Xiaowen Xu, Haiwen Hong, Zixuan Ni +4 more
The paper introduces Relay On-Policy Distillation (Relay-OPD), a method for on-policy distillation that constructs relay trajectories to address prefix failure and improve performance.