Abstract

On-policy distillation (OPD) supervises student-generated trajectories with token-level teacher signals. Its sampled-token variant avoids the cost of full-vocabulary probabilities. Yet we find that training on fewer tokens can outperform full-token OPD, challenging the intuition that more supervision improves learning. This motivates selecting tokens by learning value. Existing disagreement-based criteria ignore probability scale: tokens assigned negligible probability by both models, termed low-low tokens, can receive large log-ratio rewards and hinder learning. We propose DIAL-OPD, a token-selection method that bridges log-probability and probability spaces by weighting reward magnitude with the logarithmic mean of teacher and student probabilities. A parameter beta controls this weighting, and the highest-scoring tokens are retained. Across 4 teacher-student pairs and 7 mathematical reasoning benchmarks, we compare DIAL-OPD with 9 baselines. Retaining only 40% of tokens, it outperforms Vanilla OPD and its full-token variants, with mean accuracy gains reaching 5.25 percentage points over Vanilla OPD, and doubles AIME25 Pass@16 from 13.33% to 26.67%. It also achieves up to an 18% relative improvement in mean accuracy over the strongest token-selection baseline at matched retention ratios. With a 4B teacher, DIAL-OPD surpasses the strongest full-token baseline using an 8B teacher at both student scales, showing that effective supervision allocation can outweigh teacher scaling. Further analysis shows that moderate beta balances suppressing low-low tokens against preserving useful disagreements. Token-level evidence reveals that DIAL-OPD filters high-reward tokens with limited reasoning value while preserving supervision critical to reasoning correctness.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Zhao, A., Xin, H., Tong, J., Fan, Y., Lu, X., Nie, P., Li, W., & Shen, X. (2026). DIAL-OPD: Learning More from Fewer Tokens in On-Policy Distillation. https://omanscience.com/en/articles/dial-opd-learning-more-from-fewer-tokens-in-on-policy-distillation

MLA 9

Zhao, Anhao, et al. "DIAL-OPD: Learning More from Fewer Tokens in On-Policy Distillation." https://omanscience.com/en/articles/dial-opd-learning-more-from-fewer-tokens-in-on-policy-distillation.

Chicago (author–date)

Zhao, Anhao, Haoran Xin, Junlong Tong, Yingqi Fan, Xuan Lu, Ping Nie, Wenjie Li, and Xiaoyu Shen. 2026. "DIAL-OPD: Learning More from Fewer Tokens in On-Policy Distillation." https://omanscience.com/en/articles/dial-opd-learning-more-from-fewer-tokens-in-on-policy-distillation.

Harvard

Zhao, A., Xin, H., Tong, J., Fan, Y., Lu, X., Nie, P., Li, W. and Shen, X. (2026) 'DIAL-OPD: Learning More from Fewer Tokens in On-Policy Distillation', Available at: https://omanscience.com/en/articles/dial-opd-learning-more-from-fewer-tokens-in-on-policy-distillation.

Vancouver

Zhao A, Xin H, Tong J, Fan Y, Lu X, Nie P, et al. DIAL-OPD: Learning More from Fewer Tokens in On-Policy Distillation. https://omanscience.com/en/articles/dial-opd-learning-more-from-fewer-tokens-in-on-policy-distillation

IEEE

A. Zhao, H. Xin, J. Tong, Y. Fan, X. Lu, P. Nie, W. Li, and X. Shen, "DIAL-OPD: Learning More from Fewer Tokens in On-Policy Distillation," https://omanscience.com/en/articles/dial-opd-learning-more-from-fewer-tokens-in-on-policy-distillation.