الباحثون

Ali Hatamizadeh

المنشورات 1

نسخة أولية وصول مفتوح

PivotOPD: Learning to Recover from Pivotal Mistakes in Multi-Turn Agents

Yinghui He, Yapei Chang, Khushi Bhardwaj وآخرون · 2026

On-policy distillation (OPD) is a promising approach for training language agents, providing dense teacher supervision on student-generated trajectories. However, in multi-turn interaction, an incorrect action changes the states the student encounters later, so errors compound across turns. In preliminary experiments a …

المؤلفون المشاركون