Authors

Youran Qi

Publications 1

Preprint Open access

Lexicographic Multi-Objective On-Policy Distillation

Reinforcement learning from verifiable rewards (RLVR) usually optimizes answer correctness, yet useful language-model behavior also requires high-quality reasoning and concise responses. Existing multi-reward post-training methods typically scalarize rewards or combine specialists without explicitly protecting a reward …

Co-authors