Abstract

On-policy distillation (OPD) enables effective capability transfer between language models, yet the mechanisms underlying its failures are not fully understood. Across code generation and mathematical reasoning, OPD with larger-scale teachers exhibits early loss plateaus, with an average final loss reduction of 25.1% after 200 updates, compared with 96.2% for self-RL teachers, obtained by further reinforcement learning (RL) training of the initial student. To understand this difference, we analyze OPD as an idealized continuous-time dynamical system in the small-learning-rate limit. Our training-log diagnostics associate these plateaus with an early decline in a gradient-based learning-signal proxy while substantial loss remains; these measurements do not establish why the underlying gradient weakens. We further prove a local recovery guarantee for teachers sufficiently close to the initial student in a shared parameterization under regularity conditions, offering a conditional explanation for the success of self-RL teachers in our experiments. Across runs with and without loss plateaus, we observe small relative parameter changes (0.025-0.098%) and high similarity between the student's representations before and after OPD (linear CKA $>0.98$ across layers). These observations suggest that limited representation adaptation may contribute to learning-signal collapse, a hypothesis that remains to be tested. Code is available at https://github.com/leizhao7/opd-learning-signals.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Zhao, L., Zhao, Q., Zuo, B., & Zhan, Q. (2026). Why On-Policy Distillation Sometimes Fails: Vanishing Learning Signals. https://omanscience.com/en/articles/why-on-policy-distillation-sometimes-fails-vanishing-learning-signals

MLA 9

Zhao, Lei, et al. "Why On-Policy Distillation Sometimes Fails: Vanishing Learning Signals." https://omanscience.com/en/articles/why-on-policy-distillation-sometimes-fails-vanishing-learning-signals.

Chicago (author–date)

Zhao, Lei, Qichao Zhao, Bowen Zuo, and Qishi Zhan. 2026. "Why On-Policy Distillation Sometimes Fails: Vanishing Learning Signals." https://omanscience.com/en/articles/why-on-policy-distillation-sometimes-fails-vanishing-learning-signals.

Harvard

Zhao, L., Zhao, Q., Zuo, B. and Zhan, Q. (2026) 'Why On-Policy Distillation Sometimes Fails: Vanishing Learning Signals', Available at: https://omanscience.com/en/articles/why-on-policy-distillation-sometimes-fails-vanishing-learning-signals.

Vancouver

Zhao L, Zhao Q, Zuo B, Zhan Q. Why On-Policy Distillation Sometimes Fails: Vanishing Learning Signals. https://omanscience.com/en/articles/why-on-policy-distillation-sometimes-fails-vanishing-learning-signals

IEEE

L. Zhao, Q. Zhao, B. Zuo, and Q. Zhan, "Why On-Policy Distillation Sometimes Fails: Vanishing Learning Signals," https://omanscience.com/en/articles/why-on-policy-distillation-sometimes-fails-vanishing-learning-signals.