الباحثون

William Hoy

المنشورات 1

نسخة أولية وصول مفتوح

Asynchronous Is Nearly Free for Evolution Strategies on Long-Horizon Agentic Tasks

William Hoy, Jingxuan Fan, Nurcin Celik وآخرون · 2026

LLM-based long-horizon agentic post-training is often bottlenecked by rollout generation: trajectories span many interaction turns, completion times vary substantially, and synchronous update barriers leave faster workers waiting for stragglers. Asynchronous reinforcement learning which has been adopted in LLM post-tra …

المؤلفون المشاركون