الباحثون

Hangoo Kang

المنشورات 2

نسخة أولية وصول مفتوح

Sharpening Tax in Post-Training

Changdae Oh, Qi Zeng, Qi Qi وآخرون · 2026

An emerging hypothesis about reinforcement learning (RL) post-training of large language models (LLMs) is that it merely sharpens existing behaviors of a base model, improving single-shot accuracy at the cost of solution coverage. Although this trade-off has been observed in math and coding tasks, it need not extend to …

المؤلفون المشاركون