Abstract

Role prompting elicits specialized behavior from large language models through an expert identity, offering a lightweight way to guide reasoning on demanding tasks. However, evaluating or distilling complete role-prompted answers can miss useful next-token preferences when the sampled solution remains incorrect. Transferring these preferences also requires an objective that reaches alternatives the student rarely predicts. We introduce OPSRD, which uses a fixed expert role as privileged teaching context for on-policy self-distillation without reference solutions. A role-free student generates a trajectory, and a frozen instance of the same base model supplies role-conditioned distributions on its exact prefixes, exposing alternatives beyond the sampled continuation. Teacher-weighted forward KL targets alternatives the student underestimates, with clipping to limit individual vocabulary contributions. Supervision is restricted to the highest-entropy half of student positions, concentrating learning where predictions are uncertain. Experiments on three competition-math benchmarks with Qwen3-1.7B, 4B, and 8B show improvements over the base models without role prompts at inference. Forward KL achieves the highest macro-averaged accuracy among the three evaluated divergences at every scale. Code is available at https://github.com/zhansan114514/OPSRD.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Ren, W., Zhang, Y., Li, H., Qi, Z., Zhang, H., & Wang, N. (2026). OPSRD: On-Policy Self-Role Distillation. https://omanscience.com/en/articles/opsrd-on-policy-self-role-distillation

MLA 9

Ren, Weijie, et al. "OPSRD: On-Policy Self-Role Distillation." https://omanscience.com/en/articles/opsrd-on-policy-self-role-distillation.

Chicago (author–date)

Ren, Weijie, Yanwen Zhang, Hao Li, Zhuolin Qi, Hengyi Zhang, and Naibo Wang. 2026. "OPSRD: On-Policy Self-Role Distillation." https://omanscience.com/en/articles/opsrd-on-policy-self-role-distillation.

Harvard

Ren, W., Zhang, Y., Li, H., Qi, Z., Zhang, H. and Wang, N. (2026) 'OPSRD: On-Policy Self-Role Distillation', Available at: https://omanscience.com/en/articles/opsrd-on-policy-self-role-distillation.

Vancouver

Ren W, Zhang Y, Li H, Qi Z, Zhang H, Wang N. OPSRD: On-Policy Self-Role Distillation. https://omanscience.com/en/articles/opsrd-on-policy-self-role-distillation

IEEE

W. Ren, Y. Zhang, H. Li, Z. Qi, H. Zhang, and N. Wang, "OPSRD: On-Policy Self-Role Distillation," https://omanscience.com/en/articles/opsrd-on-policy-self-role-distillation.