Abstract

On-policy distillation (OPD) trains a student on its own generated responses using token-level teacher supervision. However, uniform weighting overlooks differences in token learning value, while existing weighting methods rely on predefined mappings from prediction signals to token weights. These mappings are not learned from the effectiveness of the resulting student updates, limiting their ability to adapt to evolving learning needs. In this paper, we propose MetaOPD, a bilevel optimization framework that jointly learns the student model and a lightweight token-weighting network. The inner objective updates the student through weighted OPD, while the outer objective optimizes the weighting network using validation loss on reference solutions after a virtual student update. Differentiating through this update connects weighting decisions to their effects on post-update performance, allowing the mapping from prediction signals to token weights to evolve alongside the student. Experiments on six mathematical reasoning and three out-of-domain datasets, covering two student scales and seven baselines, demonstrate the effectiveness of MetaOPD, with Avg@8/Pass@8 gains over OPD of 1.99/5.97 percentage points for the 0.6B student and 2.25/6.41 points for the 1.7B student.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Wang, Z., Dong, X., Wang, Y., Lu, P., Wei, X., Kuang, K., Wu, F., Dai, Z., & Zhang, M. (2026). MetaOPD: Meta-Learned Token Weighting for On-Policy Distillation. https://omanscience.com/en/articles/metaopd-meta-learned-token-weighting-for-on-policy-distillation

MLA 9

Wang, Zipeng, et al. "MetaOPD: Meta-Learned Token Weighting for On-Policy Distillation." https://omanscience.com/en/articles/metaopd-meta-learned-token-weighting-for-on-policy-distillation.

Chicago (author–date)

Wang, Zipeng, Xinpeng Dong, Yuefan Wang, Pingchen Lu, Xian Wei, Kun Kuang, Fei Wu, Zhongxiang Dai, and Min Zhang. 2026. "MetaOPD: Meta-Learned Token Weighting for On-Policy Distillation." https://omanscience.com/en/articles/metaopd-meta-learned-token-weighting-for-on-policy-distillation.

Harvard

Wang, Z., Dong, X., Wang, Y., Lu, P., Wei, X., Kuang, K., Wu, F., Dai, Z. and Zhang, M. (2026) 'MetaOPD: Meta-Learned Token Weighting for On-Policy Distillation', Available at: https://omanscience.com/en/articles/metaopd-meta-learned-token-weighting-for-on-policy-distillation.

Vancouver

Wang Z, Dong X, Wang Y, Lu P, Wei X, Kuang K, et al. MetaOPD: Meta-Learned Token Weighting for On-Policy Distillation. https://omanscience.com/en/articles/metaopd-meta-learned-token-weighting-for-on-policy-distillation

IEEE

Z. Wang, X. Dong, Y. Wang, P. Lu, X. Wei, K. Kuang, F. Wu, Z. Dai, and M. Zhang, "MetaOPD: Meta-Learned Token Weighting for On-Policy Distillation," https://omanscience.com/en/articles/metaopd-meta-learned-token-weighting-for-on-policy-distillation.