Abstract

Multi-agent systems (MAS) split a task across specialized roles and are promising on complex tasks, yet a prevailing approach relies on inference-time orchestration alone. General-purpose APIs are costly and hard to customize, while small models with role prompts rarely develop stable role competence or reliable collaboration, so post-training a MAS jointly is central. Most attempts use reinforcement learning, whose team-level reward leaves undetermined which step of which agent brought about the outcome, while local rewards need redesigning per task. On-policy distillation (OPD) gives token-level teacher supervision on trajectories the student samples, a denser signal needing no local reward, yet is underexplored for the interdependent agents of a MAS. Two difficulties arise: building complementary specialization from a judgement of which role a behavior belongs to while preserving the knowledge all roles need, and turning cross-agent collaborative information into supervision OPD can exploit. We present MAS-OPD, where Role-Advantage Specialization defines the role advantage as the difference between the teacher signals under target and non-target role conditions, and Privileged Attribution for Coordination attributes an interaction conflict to its source and supplies it to the teacher alone as privileged information. Extensive experiments on code and mathematics benchmarks show that MAS-OPD attains the highest mean score at both student scales and leads the agents to develop clearer role specialization and more effective collaborative behavior.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Zhong, Q., Zheng, M., Song, M., Jiang, H., Su, J., Ji, H., Zhang, L., & Fang, J. (2026). MAS-OPD: On-Policy Distillation for Multi-agent Systems. https://omanscience.com/en/articles/mas-opd-on-policy-distillation-for-multi-agent-systems

MLA 9

Zhong, Qiyong, et al. "MAS-OPD: On-Policy Distillation for Multi-agent Systems." https://omanscience.com/en/articles/mas-opd-on-policy-distillation-for-multi-agent-systems.

Chicago (author–date)

Zhong, Qiyong, Mao Zheng, Mingyang Song, Houcheng Jiang, Jiajie Su, Huwei Ji, Li Zhang, and Junfeng Fang. 2026. "MAS-OPD: On-Policy Distillation for Multi-agent Systems." https://omanscience.com/en/articles/mas-opd-on-policy-distillation-for-multi-agent-systems.

Harvard

Zhong, Q., Zheng, M., Song, M., Jiang, H., Su, J., Ji, H., Zhang, L. and Fang, J. (2026) 'MAS-OPD: On-Policy Distillation for Multi-agent Systems', Available at: https://omanscience.com/en/articles/mas-opd-on-policy-distillation-for-multi-agent-systems.

Vancouver

Zhong Q, Zheng M, Song M, Jiang H, Su J, Ji H, et al. MAS-OPD: On-Policy Distillation for Multi-agent Systems. https://omanscience.com/en/articles/mas-opd-on-policy-distillation-for-multi-agent-systems

IEEE

Q. Zhong, M. Zheng, M. Song, H. Jiang, J. Su, H. Ji, L. Zhang, and J. Fang, "MAS-OPD: On-Policy Distillation for Multi-agent Systems," https://omanscience.com/en/articles/mas-opd-on-policy-distillation-for-multi-agent-systems.