Abstract

Modern agentic systems combine an AI model with a harness that controls execution and environmental interactions. Harness design strongly affects long-horizon performance, yet its combinatorial search space demands substantial human effort that must be repeated as models change. Existing automated methods explore this space narrowly, optimizing only components such as prompts or skills or becoming trapped by fixed, exploitative search strategies. We introduce MILO (Meta-evolutionary Island Orchestration), a framework that co-evolves agent harnesses and the strategy used to discover them. MILO combines: (i) hierarchical lineage memory over island-based trees, using rejected mutations as negative evidence; (ii) per-island mutator agents that rewrite complete harnesses using global search history and parent-specific feedback; and (iii) an orchestrator that adapts search through lineage grafting and speciation, mutator reassignment and curriculum revision. Across Terminal-Bench 2.1, PaperBench, and DeepSWE, MILO-discovered harnesses outperform eight state-of-the-art harnesses and six search methods using frontier (Opus 4.8) and open-weight (gpt-oss-120b) models. With Opus 4.8, MILO improves resolution over its initial harness by $+12.0\%$, $+28.3\%$, and $+10.3\%$, respectively, compared with best prior-search gains of $+4.5\%$, $+18.3\%$, and $0\%$. On Terminal-Bench 2.1, it achieves $86.1 \pm 2.0\%$, exceeding the official leaderboard's top entry ($83.8 \pm 2.3\%$) while using 26\% fewer tokens than its initial harness. On EinsteinArena open problems, MILO improves best-known upper bounds for Erdős minimum-overlap ($0.3808586 \to 0.3808568$) and the first and third autocorrelation inequalities ($1.50274365 \to 1.50274360$; $1.45081 \to 1.44889$).

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Jana, P., Goswami, M., Liu, H., Li, X., Huang, L., Huang, Z., Huang, Z., Blöbaum, P., Deoras, A., Jain, P., Kanakaris, N., & Genc, S. (2026). MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution. https://omanscience.com/en/articles/milo-automated-harness-discovery-via-orchestrated-multi-agent-evolution

MLA 9

Jana, Prithwish, et al. "MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution." https://omanscience.com/en/articles/milo-automated-harness-discovery-via-orchestrated-multi-agent-evolution.

Chicago (author–date)

Jana, Prithwish, Mononito Goswami, Hao Liu, Xinyu Li, Langlin Huang, Zhehui Huang, Zhishen Huang, Patrick Blöbaum, Anoop Deoras, Purak Jain, Nikos Kanakaris, and Sahika Genc. 2026. "MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution." https://omanscience.com/en/articles/milo-automated-harness-discovery-via-orchestrated-multi-agent-evolution.

Harvard

Jana, P., Goswami, M., Liu, H., Li, X., Huang, L., Huang, Z., Huang, Z., Blöbaum, P., Deoras, A., Jain, P., Kanakaris, N. and Genc, S. (2026) 'MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution', Available at: https://omanscience.com/en/articles/milo-automated-harness-discovery-via-orchestrated-multi-agent-evolution.

Vancouver

Jana P, Goswami M, Liu H, Li X, Huang L, Huang Z, et al. MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution. https://omanscience.com/en/articles/milo-automated-harness-discovery-via-orchestrated-multi-agent-evolution

IEEE

P. Jana, M. Goswami, H. Liu, X. Li, L. Huang, Z. Huang, Z. Huang, P. Blöbaum, A. Deoras, P. Jain, N. Kanakaris, and S. Genc, "MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution," https://omanscience.com/en/articles/milo-automated-harness-discovery-via-orchestrated-multi-agent-evolution.