الباحثون

TanLin Li

المنشورات 1

نسخة أولية وصول مفتوح

Spend Teacher Tokens Where They Matter: Success-Referenced On-Policy Distillation

Xiang Chen, Futao Su, Kong Wang وآخرون · 2026

On-policy distillation (OPD) combines student-generated rollouts with dense token-level supervision from a teacher, but providing such supervision for every rollout requires substantial teacher computation. We introduce Success-Referenced On-Policy Distillation (SR-OPD), which reduces this cost by selecting which promp …

المؤلفون المشاركون