Preprint Open access
Learning to Revise Reasoning with Segment-wise On-Policy Distillation
On-policy distillation (OPD) improves large language model reasoning by training students on their own rollouts with dense token-wise supervision from the teacher. However, token-wise OPD does not explicitly provide a coherent alternative reasoning step showing how the student's step could be revised to improve subsequ …