Abstract

Repository-level software engineering (SWE) comprises heterogeneous task categories, whose progress under pooled agentic reinforcement learning can be uneven: gains in some categories coincide with regressions in others, while aggregate resolution obscures these changes. Motivated by this category see-saw, we develop a category-aware expert-training and policy-integration framework. Executable task construction and SWE Labeler, an evidence-grounded multi-axis labeling system, organize the training pools. Initial category-specific RL improves average training success while leaving uneven instance-level progress, motivating explicit consolidation of successful behavior and policy-adaptive task selection. Same-origin category experts alternate long-horizon Agentic-miniRL with Refresh-Repair-Expand (RRE): the updated policy refreshes instance mastery, reuses its own verified successful trajectories for Repair SFT, and reselects tasks for further RL. Label-routed multi-teacher on-policy distillation (MOPD) consolidates the experts into one deployable student, with ReLU-gated reward extrapolation keeping only each teacher's improving direction over the reference. Expert training and policy integration require no external model to provide solution trajectories or action targets. We evaluate Pooled RL and Balanced RL, expert development, and single-model integration through aggregate and per-category resolution, the minimum category lift over each joint-RL baseline, and expert-gain recovery. The final MOPD policy achieves mean resolution of 58.04% on Pro-618 and 59.00% on SWE-bench Multilingual, improving over the base model by 5.39 and 2.78 percentage points, respectively.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Zhao, J., Jiang, Z., Zheng, S., Shan, M., Xu, X., & Qu, L. (2026). One to More, More to One: Category-Aware Iterative Expert Training for Software Engineering Agents. https://omanscience.com/en/articles/one-to-more-more-to-one-category-aware-iterative-expert-training-for-software-engineering-agents

MLA 9

Zhao, Jie, et al. "One to More, More to One: Category-Aware Iterative Expert Training for Software Engineering Agents." https://omanscience.com/en/articles/one-to-more-more-to-one-category-aware-iterative-expert-training-for-software-engineering-agents.

Chicago (author–date)

Zhao, Jie, Ziyu Jiang, Suhang Zheng, Minghui Shan, Xiaoxiao Xu, and Lin Qu. 2026. "One to More, More to One: Category-Aware Iterative Expert Training for Software Engineering Agents." https://omanscience.com/en/articles/one-to-more-more-to-one-category-aware-iterative-expert-training-for-software-engineering-agents.

Harvard

Zhao, J., Jiang, Z., Zheng, S., Shan, M., Xu, X. and Qu, L. (2026) 'One to More, More to One: Category-Aware Iterative Expert Training for Software Engineering Agents', Available at: https://omanscience.com/en/articles/one-to-more-more-to-one-category-aware-iterative-expert-training-for-software-engineering-agents.

Vancouver

Zhao J, Jiang Z, Zheng S, Shan M, Xu X, Qu L. One to More, More to One: Category-Aware Iterative Expert Training for Software Engineering Agents. https://omanscience.com/en/articles/one-to-more-more-to-one-category-aware-iterative-expert-training-for-software-engineering-agents

IEEE

J. Zhao, Z. Jiang, S. Zheng, M. Shan, X. Xu, and L. Qu, "One to More, More to One: Category-Aware Iterative Expert Training for Software Engineering Agents," https://omanscience.com/en/articles/one-to-more-more-to-one-category-aware-iterative-expert-training-for-software-engineering-agents.