Abstract

On-policy self-distillation (OPSD) uses reference solutions as privileged hindsight to supervise student-generated reasoning trajectories. However, reference-based guidance may explain a correct solution without addressing why the student's own reasoning fails. This reasoning mismatch between the guidance provided and the correction needed can encourage the student to borrow correct conclusions while leaving its reasoning errors unresolved. Moreover, applying the same hindsight throughout the trajectory risks a distillation trap, where unnecessary constraints on valid reasoning compete with correction of substantive errors. To address these issues, we propose Root-Cause-Guided On-Policy Distillation (RC-OPD), which uses repairs of the student's own reasoning to provide guidance that addresses its specific errors while building on valid progress. For each failed attempt, RC-OPD locates the earliest substantive error, develops a local correction, and uses the corrected intermediate result as an anchor for the valid prefix. An iterative diagnosis--repair--continuation process tests the repairs through student continuation, identifying further errors within a fixed repair budget. For repair chains that reach a correct answer, root--cause--guided distillation uses failure diagnoses and corrective goals to supervise the erroneous segments, while anchor-guided distillation supports the corresponding valid prefixes with reasoning chains leading to the repaired intermediate results. We evaluate RC-OPD across multiple datasets and model scales. Extensive experiments and analyses show that it mitigates reasoning mismatch and the distillation trap, yielding substantial performance gains.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Shen, C., Yao, H., Yu, W., Jin, S., Zhang, X., & Xu, J. (2026). Learning from Repaired Reasoning: Root-Cause-Guided On-Policy Distillation. https://omanscience.com/en/articles/learning-from-repaired-reasoning-root-cause-guided-on-policy-distillation

MLA 9

Shen, Chenglei, et al. "Learning from Repaired Reasoning: Root-Cause-Guided On-Policy Distillation." https://omanscience.com/en/articles/learning-from-repaired-reasoning-root-cause-guided-on-policy-distillation.

Chicago (author–date)

Shen, Chenglei, Haoyang Yao, Weijie Yu, Song Jin, Xiao Zhang, and Jun Xu. 2026. "Learning from Repaired Reasoning: Root-Cause-Guided On-Policy Distillation." https://omanscience.com/en/articles/learning-from-repaired-reasoning-root-cause-guided-on-policy-distillation.

Harvard

Shen, C., Yao, H., Yu, W., Jin, S., Zhang, X. and Xu, J. (2026) 'Learning from Repaired Reasoning: Root-Cause-Guided On-Policy Distillation', Available at: https://omanscience.com/en/articles/learning-from-repaired-reasoning-root-cause-guided-on-policy-distillation.

Vancouver

Shen C, Yao H, Yu W, Jin S, Zhang X, Xu J. Learning from Repaired Reasoning: Root-Cause-Guided On-Policy Distillation. https://omanscience.com/en/articles/learning-from-repaired-reasoning-root-cause-guided-on-policy-distillation

IEEE

C. Shen, H. Yao, W. Yu, S. Jin, X. Zhang, and J. Xu, "Learning from Repaired Reasoning: Root-Cause-Guided On-Policy Distillation," https://omanscience.com/en/articles/learning-from-repaired-reasoning-root-cause-guided-on-policy-distillation.