Preprint Open access
PivotOPD: Learning to Recover from Pivotal Mistakes in Multi-Turn Agents
On-policy distillation (OPD) is a promising approach for training language agents, providing dense teacher supervision on student-generated trajectories. However, in multi-turn interaction, an incorrect action changes the states the student encounters later, so errors compound across turns. In preliminary experiments a …