الملخص
On-policy distillation (OPD) efficiently transfers capabilities from a stronger teacher to a student language model through dense token-level supervision on student-generated rollouts and has shown promise on complex tasks such as mathematical reasoning. However, in multi-turn agentic tasks, student decisions shape subsequent observations, causing early errors to accumulate across turns. The resulting trajectories can drift away from the teacher's rollout distribution, making the teacher's token-level supervision less reliable or even counterproductive for OPD training. To address this issue, we introduce STI-OPD, a stochastic teacher intervention framework for multi-turn agentic OPD. During multi-turn interaction, STI-OPD uses teacher intervention guided by teacher-student policy discrepancy to replace the student's proposed action with a teacher-generated one to maximize the acquisition of reliable supervision. We further develop a stochastic intervention strategy, addressing the limitations of previous threshold-based or fixed-schedule approaches, that estimates policy discrepancy using KL divergence and maps it to an intervention probability. By sampling whether to intervene from this probability, STI-OPD adaptively balances teacher control with student exploration. To learn from the resulting mixed-policy trajectories, we introduce an Importance-Weighted Reverse KL objective that corrects the token sampling mismatch between teacher-generated responses and the student policy to preserve the original OPD objective. Across tool-integrated reasoning and long-horizon interaction, STI-OPD outperforms the strongest prior OPD baseline on every evaluated benchmark and student size. Ablations further show that both discrepancy-guided intervention and importance weighting contribute to these gains.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Liu, J., Luo, L., Chen, Z., Mao, Q., Vu, T. T., & Haffari, G. (2026). Stochastic Teacher Intervention for Agentic On-Policy Distillation. https://omanscience.com/ar/articles/stochastic-teacher-intervention-for-agentic-on-policy-distillation
MLA 9
Liu, Junnan, et al. "Stochastic Teacher Intervention for Agentic On-Policy Distillation." https://omanscience.com/ar/articles/stochastic-teacher-intervention-for-agentic-on-policy-distillation.
شيكاغو (المؤلف–التاريخ)
Liu, Junnan, Linhao Luo, Zhijun Chen, Qianren Mao, Thuy-Trang Vu, and Gholamreza Haffari. 2026. "Stochastic Teacher Intervention for Agentic On-Policy Distillation." https://omanscience.com/ar/articles/stochastic-teacher-intervention-for-agentic-on-policy-distillation.
هارفارد
Liu, J., Luo, L., Chen, Z., Mao, Q., Vu, T. T. and Haffari, G. (2026) 'Stochastic Teacher Intervention for Agentic On-Policy Distillation', Available at: https://omanscience.com/ar/articles/stochastic-teacher-intervention-for-agentic-on-policy-distillation.
فانكوفر
Liu J, Luo L, Chen Z, Mao Q, Vu TT, Haffari G. Stochastic Teacher Intervention for Agentic On-Policy Distillation. https://omanscience.com/ar/articles/stochastic-teacher-intervention-for-agentic-on-policy-distillation
IEEE
J. Liu, L. Luo, Z. Chen, Q. Mao, T. T. Vu, and G. Haffari, "Stochastic Teacher Intervention for Agentic On-Policy Distillation," https://omanscience.com/ar/articles/stochastic-teacher-intervention-for-agentic-on-policy-distillation.