الباحثون

Zhenya Huang

المنشورات 2

نسخة أولية وصول مفتوح

Learning from Evolving Errors: Adaptive Iterative Repair for On-Policy Distillation

Rui Li, Liyang He, Zheng Zhang وآخرون · 2026

On-policy self-distillation (OPSD) supplies dense token-level feedback on trajectories sampled from the student's own policy, a richer training signal than the outcome-level rewards of reinforcement learning. This feedback comes from a teacher conditioned on a full reference solution unavailable to the student. The ref …

المؤلفون المشاركون