Abstract

Reasoning distillation from powerful teacher models to smaller students faces the Gap Curse: as teachers grow more sophisticated, their complex distributions increasingly diverge from what students can approximate, causing performance degradation. Existing mitigation strategies either filter out challenging examples through data selection or introduce weaker intermediate assistant models, inherently compromising supervision coverage or quality. We propose Teacher Alignment, which directly adapts the teacher toward the student's distribution without discarding data or degrading reasoning quality. However, naive alignment through standard knowledge distillation triggers catastrophic collapse of the teacher's reasoning capabilities. To address this, we reformulate teacher alignment as reinforcement learning and introduce TeacherGRPO, built on Group Relative Policy Optimization with two key innovations: (i) Curriculum Selective Alignment applies dual token- and distribution-level curricula to focus rewards on high-signal reasoning gaps while filtering noise from trivial tokens and uncertain tail distributions, and (ii) Importance-Adaptive Length Regularization selectively penalizes verbose redundancy while preserving pedagogically critical reasoning steps. The aligned teacher then distills knowledge to students via standard pipelines. Extensive experiments show TeacherGRPO significantly outperforms baselines across diverse reasoning benchmarks and distillation methods. Our code is available at https://github.com/LzyFischer/TeacherGRPO.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Lei, Z., Chen, Z., Zhu, Y., Feng, S., Zheng, Z., Guo, R., Dong, Y., & Li, J. (2026). TeacherGRPO: Closing the Capacity Gap in Reasoning Distillation via Teacher Alignment. https://omanscience.com/en/articles/teachergrpo-closing-the-capacity-gap-in-reasoning-distillation-via-teacher-alignment

MLA 9

Lei, Zhenyu, et al. "TeacherGRPO: Closing the Capacity Gap in Reasoning Distillation via Teacher Alignment." https://omanscience.com/en/articles/teachergrpo-closing-the-capacity-gap-in-reasoning-distillation-via-teacher-alignment.

Chicago (author–date)

Lei, Zhenyu, Zihan Chen, Yaochen Zhu, Shangbin Feng, Zaiyi Zheng, Ruocheng Guo, Yushun Dong, and Jundong Li. 2026. "TeacherGRPO: Closing the Capacity Gap in Reasoning Distillation via Teacher Alignment." https://omanscience.com/en/articles/teachergrpo-closing-the-capacity-gap-in-reasoning-distillation-via-teacher-alignment.

Harvard

Lei, Z., Chen, Z., Zhu, Y., Feng, S., Zheng, Z., Guo, R., Dong, Y. and Li, J. (2026) 'TeacherGRPO: Closing the Capacity Gap in Reasoning Distillation via Teacher Alignment', Available at: https://omanscience.com/en/articles/teachergrpo-closing-the-capacity-gap-in-reasoning-distillation-via-teacher-alignment.

Vancouver

Lei Z, Chen Z, Zhu Y, Feng S, Zheng Z, Guo R, et al. TeacherGRPO: Closing the Capacity Gap in Reasoning Distillation via Teacher Alignment. https://omanscience.com/en/articles/teachergrpo-closing-the-capacity-gap-in-reasoning-distillation-via-teacher-alignment

IEEE

Z. Lei, Z. Chen, Y. Zhu, S. Feng, Z. Zheng, R. Guo, Y. Dong, and J. Li, "TeacherGRPO: Closing the Capacity Gap in Reasoning Distillation via Teacher Alignment," https://omanscience.com/en/articles/teachergrpo-closing-the-capacity-gap-in-reasoning-distillation-via-teacher-alignment.