Authors

Xialiang Tong

Publications 3

Preprint Open access

Outcome-Guided On-Policy Self-Distillation

On-policy self-distillation (OPSD) provides denser token-level supervision and better computational efficiency than Reinforcement Learning with Verifiable Rewards (RLVR). However, this denser supervision may introduce substantial noise and training instability. Existing improvements often rely on high-variance per-toke …

Co-authors