Abstract
Self-evolving LLM agents promise to improve autonomously through continual interaction and learning, reducing their dependence on manually curated supervision. Realizing this promise requires not only updating the agent, but also evolving its training experience as its capabilities change. However, most existing pipelines rely on static datasets or separately updated synthesis models, causing previously useful tasks to become trivial while overly difficult tasks remain uninformative. This growing mismatch between agent capability and training experience limits sustained self-improvement. To address this problem, we propose SynCo, an agentic data synthesis co-training framework for self-evolving LLMs based on multi-agent reinforcement learning. SynCo jointly optimizes two independently parameterized agents: a Synthesizer that constructs training tasks from the Reasoner's evolving capability state, and a Reasoner that learns from the resulting experience. Each synthesized task induces multiple Reasoner rollouts whose outcomes provide complementary rewards to both agents. Correctness feedback improves the Reasoner, while task quality, answer reliability, and outcome-grounded teachability guide the Synthesizer. Their updates are fed back into subsequent synthesis rounds, allowing the task-solving policy and its training distribution to evolve together. Extensive experiments across eight mathematical reasoning benchmarks demonstrate that SynCo substantially outperforms a broad range of existing synthetic-data methods and controlled baselines, achieving the strongest overall performance while deriving most of its gains from previously unsolved problems.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Yang, W., Li, S., Qin, Y., Wang, Y., Wang, M., Li, S., Yang, T., Li, J., Thomason, J., Ma, X., & Zhao, Y. (2026). SynCo: Data Synthesis Co-Training for Self-Evolving LLMs via Multi-Agent Reinforcement Learning. https://omanscience.com/en/articles/synco-data-synthesis-co-training-for-self-evolving-llms-via-multi-agent-reinforcement-learning
MLA 9
Yang, Wei, et al. "SynCo: Data Synthesis Co-Training for Self-Evolving LLMs via Multi-Agent Reinforcement Learning." https://omanscience.com/en/articles/synco-data-synthesis-co-training-for-self-evolving-llms-via-multi-agent-reinforcement-learning.
Chicago (author–date)
Yang, Wei, Shawn Li, Yuehan Qin, Yawei Wang, Mingxi Wang, Shixuan Li, Tiankai Yang, Jiate Li, Jesse Thomason, Xuezhe Ma, and Yue Zhao. 2026. "SynCo: Data Synthesis Co-Training for Self-Evolving LLMs via Multi-Agent Reinforcement Learning." https://omanscience.com/en/articles/synco-data-synthesis-co-training-for-self-evolving-llms-via-multi-agent-reinforcement-learning.
Harvard
Yang, W., Li, S., Qin, Y., Wang, Y., Wang, M., Li, S., Yang, T., Li, J., Thomason, J., Ma, X. and Zhao, Y. (2026) 'SynCo: Data Synthesis Co-Training for Self-Evolving LLMs via Multi-Agent Reinforcement Learning', Available at: https://omanscience.com/en/articles/synco-data-synthesis-co-training-for-self-evolving-llms-via-multi-agent-reinforcement-learning.
Vancouver
Yang W, Li S, Qin Y, Wang Y, Wang M, Li S, et al. SynCo: Data Synthesis Co-Training for Self-Evolving LLMs via Multi-Agent Reinforcement Learning. https://omanscience.com/en/articles/synco-data-synthesis-co-training-for-self-evolving-llms-via-multi-agent-reinforcement-learning
IEEE
W. Yang, S. Li, Y. Qin, Y. Wang, M. Wang, S. Li, T. Yang, J. Li, J. Thomason, X. Ma, and Y. Zhao, "SynCo: Data Synthesis Co-Training for Self-Evolving LLMs via Multi-Agent Reinforcement Learning," https://omanscience.com/en/articles/synco-data-synthesis-co-training-for-self-evolving-llms-via-multi-agent-reinforcement-learning.