الملخص
On-policy distillation (OPD) trains student agents through teacher supervision on their own interactions with an environment. However, in asynchronous multi-turn training, arrival-order batching can allow a few early or long rollouts to dominate learner updates while other valid rollouts become stale before being used, wasting already-generated experience. To address this problem, we introduce DivOPD, a simple learner-side batch-selection method that spreads a fixed turn budget across more rollouts and, within each rollout, prioritizes turns with larger cumulative teacher-student disagreement. Turns without usable teacher feedback are excluded. The per-turn loss and optimizer remain fixed; selection only changes which student-visited turns receive training weight. For no-progress rollouts, an optional extension briefly hands control to the teacher before returning it to the student. Across six teacher-student settings on the simulated ALFWorld, ScienceWorld, and WebShop benchmarks, with 1.5B-7B students, DivOPD raises cross-setting mean peak success rate from 77.4 to 84.4 and mean success over the last five evaluations from 71.5 to 78.6. It reaches all reported setting-specific targets with geometric-mean speedups of 1.84x in training tokens and 1.87x in learner GPU time relative to vanilla OPD. Teacher intervention further raises this last-five mean to 82.4 while retaining about 1.7x learner-GPU speedup over vanilla OPD. Code will be released at https://github.com/HanyangWang0418-oss/DivOPD.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Wang, H., Liu, Z., Chen, Z., Ruan, J., Pang, C., Su, Z., Xie, W., Zeng, Z., Zeng, K., & Zhao, T. (2026). DivOPD: Spread Wide, Look Close for Asynchronous On-Policy Distillation of Multi-turn Agents. https://omanscience.com/ar/articles/divopd-spread-wide-look-close-for-asynchronous-on-policy-distillation-of-multi-turn-agents
MLA 9
Wang, Hanyang, et al. "DivOPD: Spread Wide, Look Close for Asynchronous On-Policy Distillation of Multi-turn Agents." https://omanscience.com/ar/articles/divopd-spread-wide-look-close-for-asynchronous-on-policy-distillation-of-multi-turn-agents.
شيكاغو (المؤلف–التاريخ)
Wang, Hanyang, Zeyuan Liu, Zhengyu Chen, Jingqing Ruan, Chaoxu Pang, Zhongda Su, Wulin Xie, Zhizhao Zeng, Ke Zeng, and Tianxiang Zhao. 2026. "DivOPD: Spread Wide, Look Close for Asynchronous On-Policy Distillation of Multi-turn Agents." https://omanscience.com/ar/articles/divopd-spread-wide-look-close-for-asynchronous-on-policy-distillation-of-multi-turn-agents.
هارفارد
Wang, H., Liu, Z., Chen, Z., Ruan, J., Pang, C., Su, Z., Xie, W., Zeng, Z., Zeng, K. and Zhao, T. (2026) 'DivOPD: Spread Wide, Look Close for Asynchronous On-Policy Distillation of Multi-turn Agents', Available at: https://omanscience.com/ar/articles/divopd-spread-wide-look-close-for-asynchronous-on-policy-distillation-of-multi-turn-agents.
فانكوفر
Wang H, Liu Z, Chen Z, Ruan J, Pang C, Su Z, et al. DivOPD: Spread Wide, Look Close for Asynchronous On-Policy Distillation of Multi-turn Agents. https://omanscience.com/ar/articles/divopd-spread-wide-look-close-for-asynchronous-on-policy-distillation-of-multi-turn-agents
IEEE
H. Wang, Z. Liu, Z. Chen, J. Ruan, C. Pang, Z. Su, W. Xie, Z. Zeng, K. Zeng, and T. Zhao, "DivOPD: Spread Wide, Look Close for Asynchronous On-Policy Distillation of Multi-turn Agents," https://omanscience.com/ar/articles/divopd-spread-wide-look-close-for-asynchronous-on-policy-distillation-of-multi-turn-agents.