الباحثون

Yun He

المنشورات 3

نسخة أولية وصول مفتوح

MIMESIS: Learning User Simulators as Training Environments for Interactive Agents

Hoang Phan, Dat Huynh, Andrey Zhmoginov وآخرون · 2026

Training and evaluating interactive language agents typically requires rich user interactions, yet collecting human feedback is expensive and difficult to scale. Simulated users offer a scalable alternative, but they must both resemble real user behavior and provide useful learning experiences for agents. In contrast, …

نسخة أولية وصول مفتوح

Sharpening Tax in Post-Training

Changdae Oh, Qi Zeng, Qi Qi وآخرون · 2026

An emerging hypothesis about reinforcement learning (RL) post-training of large language models (LLMs) is that it merely sharpens existing behaviors of a base model, improving single-shot accuracy at the cost of solution coverage. Although this trade-off has been observed in math and coding tasks, it need not extend to …

المؤلفون المشاركون