Authors

Yun He

Publications 3

Preprint Open access

Sharpening Tax in Post-Training

Changdae Oh, Qi Zeng, Qi Qi et al. · 2026

An emerging hypothesis about reinforcement learning (RL) post-training of large language models (LLMs) is that it merely sharpens existing behaviors of a base model, improving single-shot accuracy at the cost of solution coverage. Although this trade-off has been observed in math and coding tasks, it need not extend to …

Co-authors