Abstract

Advances in the coding capabilities of LLM agents allow them to inspect and modify their own instructions, tools, and execution procedures. Existing approaches use this ability to search for improved agents through repeated downstream evaluation, which incurs substantial costs and ties the search to the evaluated tasks. We introduce \textbf{SelfSearch}, a reward-free search procedure in which agents modify themselves using records of previous self-improvement episodes. These records capture the reasoning, tool actions, and outcomes of earlier modification attempts, providing concrete experience for improving both task solving and self-modification. Without downstream reward signals during search, SelfSearch improves population-mean success over the initial agent in all six model--benchmark settings, with individual agents gaining up to 11.2 percentage points on Terminal-Bench 2.1. On SWE-bench Multilingual, an agent improves success by \textbf{5.0} percentage points while reducing execution cost by \textbf{38.5}\% on tasks solved by both the initial and evolved agents. SelfSearch achieves competitive task success with evaluation-guided search baselines at lower search cost. With only \textbf{\$4.03} in search cost, it produces a harness that solves \textbf{82.0}\% of Terminal-Bench 2.1 tasks with DeepSeek V4 Flash under the settings of a public nine-harness comparison, matching the top-scoring harness, Codex. These results suggest that experience gained through self-modification can improve agents' downstream capabilities and efficiency.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Yang, J., Kong, I., & Jo, Y. (2026). SelfSearch: Reward-Free Search for Self-Improving Agents. https://omanscience.com/en/articles/selfsearch-reward-free-search-for-self-improving-agents

MLA 9

Yang, Jungwoo, et al. "SelfSearch: Reward-Free Search for Self-Improving Agents." https://omanscience.com/en/articles/selfsearch-reward-free-search-for-self-improving-agents.

Chicago (author–date)

Yang, Jungwoo, Injin Kong, and Yohan Jo. 2026. "SelfSearch: Reward-Free Search for Self-Improving Agents." https://omanscience.com/en/articles/selfsearch-reward-free-search-for-self-improving-agents.

Harvard

Yang, J., Kong, I. and Jo, Y. (2026) 'SelfSearch: Reward-Free Search for Self-Improving Agents', Available at: https://omanscience.com/en/articles/selfsearch-reward-free-search-for-self-improving-agents.

Vancouver

Yang J, Kong I, Jo Y. SelfSearch: Reward-Free Search for Self-Improving Agents. https://omanscience.com/en/articles/selfsearch-reward-free-search-for-self-improving-agents

IEEE

J. Yang, I. Kong, and Y. Jo, "SelfSearch: Reward-Free Search for Self-Improving Agents," https://omanscience.com/en/articles/selfsearch-reward-free-search-for-self-improving-agents.