الملخص
Multi-turn agents trained with reinforcement learning (RL) receive a single scalar reward per trajectory, which motivates self on-policy distillation (OPD) to supply dense token-level supervision from a self-teacher with privileged task skills, letting a skill-free student internalize them. This recipe, however, is undermined by two findings in agentic tasks: privileged information alone does not always make a teacher reliable, and the benefit of teacher supervision is stage-dependent. We therefore propose RetireOPD (Self-Retiring On-Policy Distillation), which first optimizes a decoupled, skill-conditioned teacher with environment rewards and then trains a skill-free student jointly with RL and OPD. Rather than following a predefined distillation schedule, RetireOPD adopts Adaptive Retirement: the student drops the teacher on its own once their discrepancy stops shrinking and it reaches a target fraction of the teacher's success rate, after which training proceeds with RL alone. Across Qwen2.5 models from 1.5B to 7B, RetireOPD improves ALFWorld success rate over RL baseline by 14.1% to 18.8% and WebShop accuracy by 11.8% to 19.0%, and surpasses its own skill-conditioned teacher in every setting.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Yu, Y., Lu, Z., Liu, Y., Pan, Y., Wang, A., Chen, Q., Yang, H., Zhang, W., Chen, Q., & Shen, Y. (2026). RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning. https://omanscience.com/ar/articles/retireopd-self-retiring-on-policy-distillation-for-agentic-reinforcement-learning
MLA 9
Yu, Yan, et al. "RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning." https://omanscience.com/ar/articles/retireopd-self-retiring-on-policy-distillation-for-agentic-reinforcement-learning.
شيكاغو (المؤلف–التاريخ)
Yu, Yan, Zhengxi Lu, Yizhou Liu, Yichen Pan, Aozhe Wang, Qipeng Chen, Hua Yang, Wenqi Zhang, Qianglong Chen, and Yongliang Shen. 2026. "RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning." https://omanscience.com/ar/articles/retireopd-self-retiring-on-policy-distillation-for-agentic-reinforcement-learning.
هارفارد
Yu, Y., Lu, Z., Liu, Y., Pan, Y., Wang, A., Chen, Q., Yang, H., Zhang, W., Chen, Q. and Shen, Y. (2026) 'RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning', Available at: https://omanscience.com/ar/articles/retireopd-self-retiring-on-policy-distillation-for-agentic-reinforcement-learning.
فانكوفر
Yu Y, Lu Z, Liu Y, Pan Y, Wang A, Chen Q, et al. RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning. https://omanscience.com/ar/articles/retireopd-self-retiring-on-policy-distillation-for-agentic-reinforcement-learning
IEEE
Y. Yu, Z. Lu, Y. Liu, Y. Pan, A. Wang, Q. Chen, H. Yang, W. Zhang, Q. Chen, and Y. Shen, "RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning," https://omanscience.com/ar/articles/retireopd-self-retiring-on-policy-distillation-for-agentic-reinforcement-learning.