Preprint Open access
MetaOPD: Meta-Learned Token Weighting for On-Policy Distillation
On-policy distillation (OPD) trains a student on its own generated responses using token-level teacher supervision. However, uniform weighting overlooks differences in token learning value, while existing weighting methods rely on predefined mappings from prediction signals to token weights. These mappings are not lear …