الملخص
On-policy self-distillation (OPSD) supplies dense token-level feedback on trajectories sampled from the student's own policy, a richer training signal than the outcome-level rewards of reinforcement learning. This feedback comes from a teacher conditioned on a full reference solution unavailable to the student. The reference solution specifies the target but not how to move from the student's current error toward it, creating a solution-conditioned shortcut risk. We introduce AIR-OPD, an adaptive iterative repair framework for on-policy distillation that provides error-to-repair supervision. Given a failed response, a guidance generator synthesizes repair guidance for the current error. The student samples an on-policy retry with this guidance. If the retry remains incorrect, the generator produces new repair guidance for the newly observed error. At each round, a fixed teacher receives the guidance as privileged context and supervises the student on an error-aligned region of its latest failed response. Outcome-aware stage weighting favors early repair stages and credits stages whose immediate retry passes verification. We train AIR-OPD on the DAPO-Math-17K dataset and evaluate on AIME24, AIME25, and HMMT25, alongside out-of-distribution tests on MMLU-Pro and GPQA. We examine two guidance sources, self-guidance from the current student policy and external guidance from a larger model. For both Qwen3-4B and Qwen3-8B, AIR-OPD attains the best mathematical-reasoning averages, improving over the strongest baseline by up to 3.6 points, while preserving base-model performance on the out-of-distribution benchmarks.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Li, R., He, L., Zhang, Z., Huang, Z., Zhu, L., & Liu, Q. (2026). Learning from Evolving Errors: Adaptive Iterative Repair for On-Policy Distillation. https://omanscience.com/ar/articles/learning-from-evolving-errors-adaptive-iterative-repair-for-on-policy-distillation
MLA 9
Li, Rui, et al. "Learning from Evolving Errors: Adaptive Iterative Repair for On-Policy Distillation." https://omanscience.com/ar/articles/learning-from-evolving-errors-adaptive-iterative-repair-for-on-policy-distillation.
شيكاغو (المؤلف–التاريخ)
Li, Rui, Liyang He, Zheng Zhang, Zhenya Huang, Linbo Zhu, and Qi Liu. 2026. "Learning from Evolving Errors: Adaptive Iterative Repair for On-Policy Distillation." https://omanscience.com/ar/articles/learning-from-evolving-errors-adaptive-iterative-repair-for-on-policy-distillation.
هارفارد
Li, R., He, L., Zhang, Z., Huang, Z., Zhu, L. and Liu, Q. (2026) 'Learning from Evolving Errors: Adaptive Iterative Repair for On-Policy Distillation', Available at: https://omanscience.com/ar/articles/learning-from-evolving-errors-adaptive-iterative-repair-for-on-policy-distillation.
فانكوفر
Li R, He L, Zhang Z, Huang Z, Zhu L, Liu Q. Learning from Evolving Errors: Adaptive Iterative Repair for On-Policy Distillation. https://omanscience.com/ar/articles/learning-from-evolving-errors-adaptive-iterative-repair-for-on-policy-distillation
IEEE
R. Li, L. He, Z. Zhang, Z. Huang, L. Zhu, and Q. Liu, "Learning from Evolving Errors: Adaptive Iterative Repair for On-Policy Distillation," https://omanscience.com/ar/articles/learning-from-evolving-errors-adaptive-iterative-repair-for-on-policy-distillation.