الملخص

On-policy self-distillation (OPSD) provides dense token-level supervision without a second model: one network acts as teacher with the reference solution and as student with only the problem. We identify a specific failure mode of this recipe. During training, student token entropy rises past the teacher's and remains elevated, a pattern we call entropy overshoot. We trace it to both sides of distillation. The reference-conditioned teacher is confident along its answer-directed reasoning path, but this confidence transfers poorly to student-generated prefixes, making its supervision overly tied to answer-specific cues rather than reusable reasoning patterns; meanwhile, the forward KL used by OPSD continually diffuses the student's predictive distribution without pulling it back. We introduce E$^2$-OPSD to address both causes. Exemplar-guided teaching replaces the current answer with a retrieved solved neighboring problem, providing transferable reasoning guidance without revealing the destination and better matching student-reachable states. Entropy-aware distillation uses the student-teacher entropy gap to determine the direction and strength of each token's correction. E$^2$-OPSD improves math reasoning by up to 4.3 points in mean@16 over OPSD, while out-of-domain evaluations show gains over the corresponding base models of up to 4.9 points in mean@16 and 5.5 points in pass@8. Despite these gains, E$^2$-OPSD remains simple, requiring no additional forward passes or networks.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Liu, Y., Fang, M., Gu, X., Yao, C., Liu, M., Ma, T., Zheng, J., Yu, C., & Gao, Z. (2026). E$^2$-OPSD: Taming Entropy Overshoot in On-Policy Self-Distillation. https://omanscience.com/ar/articles/e-2-opsd-taming-entropy-overshoot-in-on-policy-self-distillation

MLA 9

Liu, Yifei, et al. "E$^2$-OPSD: Taming Entropy Overshoot in On-Policy Self-Distillation." https://omanscience.com/ar/articles/e-2-opsd-taming-entropy-overshoot-in-on-policy-self-distillation.

شيكاغو (المؤلف–التاريخ)

Liu, Yifei, Minghao Fang, Xinyu Gu, Chengkai Yao, Mengdi Liu, Tengfei Ma, Jiangbin Zheng, Chang Yu, and Zhangyang Gao. 2026. "E$^2$-OPSD: Taming Entropy Overshoot in On-Policy Self-Distillation." https://omanscience.com/ar/articles/e-2-opsd-taming-entropy-overshoot-in-on-policy-self-distillation.

هارفارد

Liu, Y., Fang, M., Gu, X., Yao, C., Liu, M., Ma, T., Zheng, J., Yu, C. and Gao, Z. (2026) 'E$^2$-OPSD: Taming Entropy Overshoot in On-Policy Self-Distillation', Available at: https://omanscience.com/ar/articles/e-2-opsd-taming-entropy-overshoot-in-on-policy-self-distillation.

فانكوفر

Liu Y, Fang M, Gu X, Yao C, Liu M, Ma T, et al. E$^2$-OPSD: Taming Entropy Overshoot in On-Policy Self-Distillation. https://omanscience.com/ar/articles/e-2-opsd-taming-entropy-overshoot-in-on-policy-self-distillation

IEEE

Y. Liu, M. Fang, X. Gu, C. Yao, M. Liu, T. Ma, J. Zheng, C. Yu, and Z. Gao, "E$^2$-OPSD: Taming Entropy Overshoot in On-Policy Self-Distillation," https://omanscience.com/ar/articles/e-2-opsd-taming-entropy-overshoot-in-on-policy-self-distillation.