الباحثون

Renjie Cao

المنشورات 1

نسخة أولية وصول مفتوح

ReTaCo: Residual-Target Control for On-Policy Distillation

Zixiang Ni, Zhuo Hu, Renjie Cao وآخرون · 2026

On-policy distillation (OPD) trains a student on its own generated prefixes with token-level teacher feedback, but transmitting or storing the teacher's full-vocabulary distribution at every token is costly. Entropy-aware OPD (EOPD) adds forward supervision to reverse KL to help the student recover plausible tokens it …

المؤلفون المشاركون