Yu-Chiang Frank Wang
2 indexed papers
Publications per year
Top categories
Frequent co-authors
Research Timeline
This paper proposes SpatialClaw, a training-free framework for spatial reasoning that enables open-ended, complex 3D/4D spatial reasoning.
This paper introduces Zone of Proximal Policy Optimization (ZPPO) for knowledge distillation, which keeps the teacher inside the student's current zone of development and outperforms off/on-policy distillation and GRPO.
Papers
Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients
Byung-Kwan Lee, Ximing Lu, Shizhe Diao, Minki Kang +7 more
This paper introduces Zone of Proximal Policy Optimization (ZPPO) for knowledge distillation, which keeps the teacher inside the student's current zone of development and outperforms off/on-policy dis…