Abstract

LLM-based long-horizon agentic post-training is often bottlenecked by rollout generation: trajectories span many interaction turns, completion times vary substantially, and synchronous update barriers leave faster workers waiting for stragglers. Asynchronous reinforcement learning which has been adopted in LLM post-training addresses this inefficiency by consuming trajectories as they arrive, but introduces policy lag and off-policy optimization. Evolution strategies (ES) offer a backpropagation-free alternative for LLM post-training, yet it relies on a larger number of rollouts and existing practices have remained largely synchronous. In this short-form paper, we introduce bounded-staleness asynchronous ES and demonstrate it on Endless Terminals benchmark using Qwen2.5-7B-Instruct. Across three evaluation seeds, natural Async-1 matches synchronous ES, achieving 25.9\% versus 25.4\% held-out success. Controlled schedules that delay 10\% of each update cohort by four or eight policy updates reduce success by only 1.6 and 3.1 percentage points, respectively, without explicit off-policy correction. GRPO performs better overall, reaching 29.0\% held-out success, but importantly our results show that ES tolerates moderate policy staleness with limited degradation, opening possibilities for future improvement of ES-based post-training with asynchronous algorithms. To the best of our knowledge, we are the first to demonstrate the effectiveness of sync and async ES on a multi-turn terminal style agentic coding task.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Hoy, W., Fan, J., Celik, N., & Pan, X. (2026). Asynchronous Is Nearly Free for Evolution Strategies on Long-Horizon Agentic Tasks. https://omanscience.com/en/articles/asynchronous-is-nearly-free-for-evolution-strategies-on-long-horizon-agentic-tasks

MLA 9

Hoy, William, et al. "Asynchronous Is Nearly Free for Evolution Strategies on Long-Horizon Agentic Tasks." https://omanscience.com/en/articles/asynchronous-is-nearly-free-for-evolution-strategies-on-long-horizon-agentic-tasks.

Chicago (author–date)

Hoy, William, Jingxuan Fan, Nurcin Celik, and Xu Pan. 2026. "Asynchronous Is Nearly Free for Evolution Strategies on Long-Horizon Agentic Tasks." https://omanscience.com/en/articles/asynchronous-is-nearly-free-for-evolution-strategies-on-long-horizon-agentic-tasks.

Harvard

Hoy, W., Fan, J., Celik, N. and Pan, X. (2026) 'Asynchronous Is Nearly Free for Evolution Strategies on Long-Horizon Agentic Tasks', Available at: https://omanscience.com/en/articles/asynchronous-is-nearly-free-for-evolution-strategies-on-long-horizon-agentic-tasks.

Vancouver

Hoy W, Fan J, Celik N, Pan X. Asynchronous Is Nearly Free for Evolution Strategies on Long-Horizon Agentic Tasks. https://omanscience.com/en/articles/asynchronous-is-nearly-free-for-evolution-strategies-on-long-horizon-agentic-tasks

IEEE

W. Hoy, J. Fan, N. Celik, and X. Pan, "Asynchronous Is Nearly Free for Evolution Strategies on Long-Horizon Agentic Tasks," https://omanscience.com/en/articles/asynchronous-is-nearly-free-for-evolution-strategies-on-long-horizon-agentic-tasks.