نسخة أولية وصول مفتوح
E$^2$-OPSD: Taming Entropy Overshoot in On-Policy Self-Distillation
On-policy self-distillation (OPSD) provides dense token-level supervision without a second model: one network acts as teacher with the reference solution and as student with only the problem. We identify a specific failure mode of this recipe. During training, student token entropy rises past the teacher's and remains …