Abstract
Reinforcement learning from verifiable rewards (RLVR) usually optimizes answer correctness, yet useful language-model behavior also requires high-quality reasoning and concise responses. Existing multi-reward post-training methods typically scalarize rewards or combine specialists without explicitly protecting a reward priority order. This is problematic when trade-offs are asymmetric: conciseness, for example, should not improve at the cost of correctness. We introduce Lexicographic Multi-Objective On-Policy Distillation (LMOPD), a multi-teacher method for integrating reward-specialized policies under explicit priorities. For each student rollout, LMOPD selects the specialist for the first objective whose gate detects a deficiency, then locally projects its centered log-policy correction to remove components that oppose higher-priority specialists. We evaluate 30B-A3B mixture-of-experts transformer models in two- and four-expert settings on three math benchmarks, measuring retained specialist gains. With two experts, LMOPD's point estimates fully retain the accuracy and reasoning-quality gains while acquiring $46.9\%$ of the conciseness gain. With four experts, it retains $\approx90\%$ of both the accuracy gain and reasoning-correctness gain, compared to only $\approx57\%$ by the next best evaluated baseline. Matched four-expertablations show that lexicographic routing outperforms random routing and that projection further strengthens both top-priority capabilities. Across both scales, LMOPD preserves the highest-priority capabilities more effectively than the existing baselines we evaluate, demonstrating the value of explicit priorities for specialist integration.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Jang, D., Campos, J. A., & Qi, Y. (2026). Lexicographic Multi-Objective On-Policy Distillation. https://omanscience.com/en/articles/lexicographic-multi-objective-on-policy-distillation
MLA 9
Jang, Doseok, et al. "Lexicographic Multi-Objective On-Policy Distillation." https://omanscience.com/en/articles/lexicographic-multi-objective-on-policy-distillation.
Chicago (author–date)
Jang, Doseok, Jon Ander Campos, and Youran Qi. 2026. "Lexicographic Multi-Objective On-Policy Distillation." https://omanscience.com/en/articles/lexicographic-multi-objective-on-policy-distillation.
Harvard
Jang, D., Campos, J. A. and Qi, Y. (2026) 'Lexicographic Multi-Objective On-Policy Distillation', Available at: https://omanscience.com/en/articles/lexicographic-multi-objective-on-policy-distillation.
Vancouver
Jang D, Campos JA, Qi Y. Lexicographic Multi-Objective On-Policy Distillation. https://omanscience.com/en/articles/lexicographic-multi-objective-on-policy-distillation
IEEE
D. Jang, J. A. Campos, and Y. Qi, "Lexicographic Multi-Objective On-Policy Distillation," https://omanscience.com/en/articles/lexicographic-multi-objective-on-policy-distillation.