Abstract

Self-evolving LLM agents promise to improve autonomously through continual interaction and learning, reducing their dependence on manually curated supervision. Realizing this promise requires not only updating the agent, but also evolving its training experience as its capabilities change. However, most existing pipelines rely on static datasets or separately updated synthesis models, causing previously useful tasks to become trivial while overly difficult tasks remain uninformative. This growing mismatch between agent capability and training experience limits sustained self-improvement. To address this problem, we propose SynCo, an agentic data synthesis co-training framework for self-evolving LLMs based on multi-agent reinforcement learning. SynCo jointly optimizes two independently parameterized agents: a Synthesizer that constructs training tasks from the Reasoner's evolving capability state, and a Reasoner that learns from the resulting experience. Each synthesized task induces multiple Reasoner rollouts whose outcomes provide complementary rewards to both agents. Correctness feedback improves the Reasoner, while task quality, answer reliability, and outcome-grounded teachability guide the Synthesizer. Their updates are fed back into subsequent synthesis rounds, allowing the task-solving policy and its training distribution to evolve together. Extensive experiments across eight mathematical reasoning benchmarks demonstrate that SynCo substantially outperforms a broad range of existing synthetic-data methods and controlled baselines, achieving the strongest overall performance while deriving most of its gains from previously unsolved problems.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Yang, W., Li, S., Qin, Y., Wang, Y., Wang, M., Li, S., Yang, T., Li, J., Thomason, J., Ma, X., & Zhao, Y. (2026). SynCo: Data Synthesis Co-Training for Self-Evolving LLMs via Multi-Agent Reinforcement Learning. https://omanscience.com/en/articles/synco-data-synthesis-co-training-for-self-evolving-llms-via-multi-agent-reinforcement-learning

MLA 9

Yang, Wei, et al. "SynCo: Data Synthesis Co-Training for Self-Evolving LLMs via Multi-Agent Reinforcement Learning." https://omanscience.com/en/articles/synco-data-synthesis-co-training-for-self-evolving-llms-via-multi-agent-reinforcement-learning.

Chicago (author–date)

Yang, Wei, Shawn Li, Yuehan Qin, Yawei Wang, Mingxi Wang, Shixuan Li, Tiankai Yang, Jiate Li, Jesse Thomason, Xuezhe Ma, and Yue Zhao. 2026. "SynCo: Data Synthesis Co-Training for Self-Evolving LLMs via Multi-Agent Reinforcement Learning." https://omanscience.com/en/articles/synco-data-synthesis-co-training-for-self-evolving-llms-via-multi-agent-reinforcement-learning.

Harvard

Yang, W., Li, S., Qin, Y., Wang, Y., Wang, M., Li, S., Yang, T., Li, J., Thomason, J., Ma, X. and Zhao, Y. (2026) 'SynCo: Data Synthesis Co-Training for Self-Evolving LLMs via Multi-Agent Reinforcement Learning', Available at: https://omanscience.com/en/articles/synco-data-synthesis-co-training-for-self-evolving-llms-via-multi-agent-reinforcement-learning.

Vancouver

Yang W, Li S, Qin Y, Wang Y, Wang M, Li S, et al. SynCo: Data Synthesis Co-Training for Self-Evolving LLMs via Multi-Agent Reinforcement Learning. https://omanscience.com/en/articles/synco-data-synthesis-co-training-for-self-evolving-llms-via-multi-agent-reinforcement-learning

IEEE

W. Yang, S. Li, Y. Qin, Y. Wang, M. Wang, S. Li, T. Yang, J. Li, J. Thomason, X. Ma, and Y. Zhao, "SynCo: Data Synthesis Co-Training for Self-Evolving LLMs via Multi-Agent Reinforcement Learning," https://omanscience.com/en/articles/synco-data-synthesis-co-training-for-self-evolving-llms-via-multi-agent-reinforcement-learning.