الباحثون

Xian Wei

المنشورات 1

نسخة أولية وصول مفتوح

MetaOPD: Meta-Learned Token Weighting for On-Policy Distillation

Zipeng Wang, Xinpeng Dong, Yuefan Wang وآخرون · 2026

On-policy distillation (OPD) trains a student on its own generated responses using token-level teacher supervision. However, uniform weighting overlooks differences in token learning value, while existing weighting methods rely on predefined mappings from prediction signals to token weights. These mappings are not lear …

المؤلفون المشاركون