الملخص
تمت ترجمة أجزاء من هذه الصفحة آلياً وقد تحتوي على أخطاء.
Reinforcement learning can substantially improve a reasoning teacher, but it is unclear which of those improvements survive when the teacher supervises a smaller on-policy student. We study this question in mathematical reasoning by comparing teacher lineages before and after GRPO, multiple student scales, direct GRPO, and several on-policy distillation objectives. The central finding is that transfer is structured rather than scalar: teacher strength alone does not make dense distillation competitive, while an RL-improved teacher creates useful but metric-dependent student gains. This motivates Transfer-Stratified On-Policy Distillation (TS-OPD), which screens training problems by the joint sampled success of the student and teacher, routes acquisition problems to gated forward KL, routes consolidation problems to gated reverse KL, and adds an entropy brake to protect sampled coverage. Across the main comparison, TS-OPD is the strongest student objective for macro average correctness with the GRPO-improved teacher, while pass@K remains more mixed. Ablations show that the gains come from routing and token gating rather than skipping problems. These results support a transfer-aware view of OPD: stronger teachers help when the supervision direction and token budget match the student's observed ability, not merely because the teacher endpoint is stronger.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Chen, X., Shao, B., Zhu, T., Wu, B., Shou, L., Wu, F., Sun, F., & Ding, W. (2026). التقطير على السياسة المقسّم بحسب النقل للمعلمين المحسّنين بالتعلم المعزز في الاستدلال. https://omanscience.com/ar/articles/transfer-stratified-on-policy-distillation-for-rl-improved-reasoning-teachers
MLA 9
Chen, Xiaoyu, et al. "التقطير على السياسة المقسّم بحسب النقل للمعلمين المحسّنين بالتعلم المعزز في الاستدلال." https://omanscience.com/ar/articles/transfer-stratified-on-policy-distillation-for-rl-improved-reasoning-teachers.
شيكاغو (المؤلف–التاريخ)
Chen, Xiaoyu, Bo Shao, Tiangang Zhu, Bintao Wu, Linjun Shou, Fengge Wu, Feng Sun, and Wenbiao Ding. 2026. "التقطير على السياسة المقسّم بحسب النقل للمعلمين المحسّنين بالتعلم المعزز في الاستدلال." https://omanscience.com/ar/articles/transfer-stratified-on-policy-distillation-for-rl-improved-reasoning-teachers.
هارفارد
Chen, X., Shao, B., Zhu, T., Wu, B., Shou, L., Wu, F., Sun, F. and Ding, W. (2026) 'التقطير على السياسة المقسّم بحسب النقل للمعلمين المحسّنين بالتعلم المعزز في الاستدلال', Available at: https://omanscience.com/ar/articles/transfer-stratified-on-policy-distillation-for-rl-improved-reasoning-teachers.
فانكوفر
Chen X, Shao B, Zhu T, Wu B, Shou L, Wu F, et al. التقطير على السياسة المقسّم بحسب النقل للمعلمين المحسّنين بالتعلم المعزز في الاستدلال. https://omanscience.com/ar/articles/transfer-stratified-on-policy-distillation-for-rl-improved-reasoning-teachers
IEEE
X. Chen, B. Shao, T. Zhu, B. Wu, L. Shou, F. Wu, F. Sun, and W. Ding, "التقطير على السياسة المقسّم بحسب النقل للمعلمين المحسّنين بالتعلم المعزز في الاستدلال," https://omanscience.com/ar/articles/transfer-stratified-on-policy-distillation-for-rl-improved-reasoning-teachers.