Preprint Open access
A Good Self-Teacher Meets the Student Where They Are: Joint On-Policy Learning and Teaching
Reinforcement Learning (RL) from outcome rewards suffers from sparse supervision, particularly on difficult, long-horizon tasks where successful trajectories are rare and costly to generate. On-Policy Distillation (OPD) offers an attractive alternative by providing dense token-level supervision from a stronger teacher …