الباحثون

Liyang He

المنشورات 1

نسخة أولية وصول مفتوح

Learning from Evolving Errors: Adaptive Iterative Repair for On-Policy Distillation

Rui Li, Liyang He, Zheng Zhang وآخرون · 2026

On-policy self-distillation (OPSD) supplies dense token-level feedback on trajectories sampled from the student's own policy, a richer training signal than the outcome-level rewards of reinforcement learning. This feedback comes from a teacher conditioned on a full reference solution unavailable to the student. The ref …

المؤلفون المشاركون