Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:
ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Home/Authors/Saurav Muralidharan

Saurav Muralidharan

1 indexed paper

Recent (6 mo)
1
With code
0
Influential cites
0
Benchmarked
0

Publications per year

1
26

Top categories

NLP×1

Frequent co-authors

Byung-Kwan Lee1×
Ximing Lu1×
Shizhe Diao1×
Minki Kang1×
Karan Sapra1×
Andrew Tao1×

Research Timeline

2026
Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients

This paper introduces Zone of Proximal Policy Optimization (ZPPO) for knowledge distillation, which keeps the teacher inside the student's current zone of development and outperforms off/on-policy distillation and GRPO.

Highlighted terms show continued research focus across papers

Papers

cs.CLEmpiricalRecentJun 16, 2026

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients

Byung-Kwan Lee, Ximing Lu, Shizhe Diao, Minki Kang +7 more

This paper introduces Zone of Proximal Policy Optimization (ZPPO) for knowledge distillation, which keeps the teacher inside the student's current zone of development and outperforms off/on-policy dis…

View →