Yongliang Shen
2 indexed papers
Publications per year
Top categories
Frequent co-authors
Research Timeline
This paper proposes Perceive-to-Reason (P2R), a framework for fine-grained visual reasoning that decouples perception from reasoning and introduces a new reinforcement learning strategy.
The paper introduces Relay On-Policy Distillation (Relay-OPD), a method for on-policy distillation that constructs relay trajectories to address prefix failure and improve performance.
Papers
Pass the Baton: Trajectory-Relayed On-Policy Distillation
Haolei Xu, Xiaowen Xu, Haiwen Hong, Zixuan Ni +4 more
The paper introduces Relay On-Policy Distillation (Relay-OPD), a method for on-policy distillation that constructs relay trajectories to address prefix failure and improve performance.