Abstract
Recursive self-improvement (RSI) of a model on non-verifiable tasks, such as open-ended research, faces a supervision bottleneck when its outputs exceed what even human experts can reliably assess, leaving the model itself (optimizee) as the best available optimizer and evaluator. However, a single model instance struggles to critique and improve its own complex reasoning under this homogeneous loop. To address this, we propose Multi-Agent Self-Supervision (MASS), an RSI method that alternates between evolutionary workflow optimization and supervised fine-tuning on self-generated trajectories. Guided by early findings that multi-agent topologies excel at complex reasoning, MASS prompts a single base model to iteratively propose, execute, and self-evaluate multi-agent workflows. Through an evolutionary search constrained by structural guardrails, the model optimizes these computational-graph-like orchestrations, discovering the most effective distinct roles and information routing for a given task. Over two MASS cycles with Qwen3.6-27B, the model achieves 1.2-1.6x higher performance per output tokens on four open-ended public benchmarks. Because the improved model subsequently acts as a better optimizer and evaluator, this alternating framework enables a continuous, recursive bootstrapping of the model's capabilities. Moreover, multi-agent traces are also more training-efficient: a student trained on them outperforms a single-agent student trained on 1.4x more training tokens. These findings suggest that jointly learning orchestration and bounded subagent execution from multi-agent trajectories can provide an effective signal for RSI.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Lee, H., Xu, J., Seely, J., Lee, D., Sojoudi, S., Zaharia, M., & Tang, Y. (2026). Recursive Self-Improvement through Multi-Agent Self-Supervision. https://omanscience.com/en/articles/recursive-self-improvement-through-multi-agent-self-supervision
MLA 9
Lee, Hyunin, et al. "Recursive Self-Improvement through Multi-Agent Self-Supervision." https://omanscience.com/en/articles/recursive-self-improvement-through-multi-agent-self-supervision.
Chicago (author–date)
Lee, Hyunin, Jinglue Xu, Jeffrey Seely, Donghyun Lee, Somayeh Sojoudi, Matei Zaharia, and Yujin Tang. 2026. "Recursive Self-Improvement through Multi-Agent Self-Supervision." https://omanscience.com/en/articles/recursive-self-improvement-through-multi-agent-self-supervision.
Harvard
Lee, H., Xu, J., Seely, J., Lee, D., Sojoudi, S., Zaharia, M. and Tang, Y. (2026) 'Recursive Self-Improvement through Multi-Agent Self-Supervision', Available at: https://omanscience.com/en/articles/recursive-self-improvement-through-multi-agent-self-supervision.
Vancouver
Lee H, Xu J, Seely J, Lee D, Sojoudi S, Zaharia M, et al. Recursive Self-Improvement through Multi-Agent Self-Supervision. https://omanscience.com/en/articles/recursive-self-improvement-through-multi-agent-self-supervision
IEEE
H. Lee, J. Xu, J. Seely, D. Lee, S. Sojoudi, M. Zaharia, and Y. Tang, "Recursive Self-Improvement through Multi-Agent Self-Supervision," https://omanscience.com/en/articles/recursive-self-improvement-through-multi-agent-self-supervision.