Authors

Ziyang Ding

Publications 2

Preprint Open access

Unbiased Top-$k$ Estimation for On-Policy Distillation

Linjian Meng, Siyuan Gan, Yuhan Li et al. · 2026

On-policy distillation (OPD) is becoming an important component of large language model (LLM) post-training for transferring the reasoning capability of a strong teacher LLM to a weaker student LLM. OPD trains the student by minimizing the reverse KL divergence between the teacher and the student via rollouts generated …

Co-authors