Karan Sapra
1 indexed paper
Recent (6 mo)
1With code
0Influential cites
0Benchmarked
0Publications per year
126
Top categories
NLP×1
Frequent co-authors
Research Timeline
2026
Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients
This paper introduces Zone of Proximal Policy Optimization (ZPPO) for knowledge distillation, which keeps the teacher inside the student's current zone of development and outperforms off/on-policy distillation and GRPO.
Highlighted terms show continued research focus across papers
Papers
cs.CLEmpiricalRecentJun 16, 2026
Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients
Byung-Kwan Lee, Ximing Lu, Shizhe Diao, Minki Kang +7 more
This paper introduces Zone of Proximal Policy Optimization (ZPPO) for knowledge distillation, which keeps the teacher inside the student's current zone of development and outperforms off/on-policy dis…
View →