الباحثون

Sharon Li

المنشورات 3

نسخة أولية وصول مفتوح

Sharpening Tax in Post-Training

Changdae Oh, Qi Zeng, Qi Qi وآخرون · 2026

An emerging hypothesis about reinforcement learning (RL) post-training of large language models (LLMs) is that it merely sharpens existing behaviors of a base model, improving single-shot accuracy at the cost of solution coverage. Although this trade-off has been observed in math and coding tasks, it need not extend to …

نسخة أولية وصول مفتوح

AIM: Agentic Idea Management for Automated Research

Frontier LLMs are increasingly used to automate scientific research through iterative search. We distinguish idea-driven search from solution-driven search and identify three core challenges: organizing evolving research ideas, selecting promising directions, and maintaining alignment between ideas and their implementa …

المؤلفون المشاركون