Abstract
On-policy distillation (OPD) trains a student on its own rollouts using dense supervision from a teacher. In multi-turn environments, a mistake at a critical decision step can redirect the subsequent rollout toward poor outcomes. We use low teacher confidence on student actions to select high-uncertainty steps for correction. In a controlled ALFWorld study, a single teacher correction at a low-confidence step improves subsequent student behavior and task success, motivating selective intervention during distillation. We propose UOPD, an uncertainty-aware intervention method for on-policy distillation. At low-uncertainty turns, UOPD executes student actions and applies the standard OPD loss. At high-uncertainty turns, it samples and executes teacher actions and trains the student to imitate them through supervised fine-tuning, which minimizes forward Kullback-Leibler divergence in expectation. UOPD utilizes adaptive uncertainty thresholds to target a scheduled intervention rate. Empirically, we evaluate UOPD across a broad range of agentic tasks, including ALFWorld, WebShop, and Search, demonstrating its superior performance over OPD methods and their variants. UOPD improves WebShop score by up to $15.8\%$ relative to standard OPD.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Zhang, W., Xu, P., Du, W., Zhang, J., & Cai, H. (2026). UOPD: Uncertainty-Aware Intervention for On-Policy Distillation of Multi-Turn Agents. https://omanscience.com/en/articles/uopd-uncertainty-aware-intervention-for-on-policy-distillation-of-multi-turn-agents
MLA 9
Zhang, Wenbo, et al. "UOPD: Uncertainty-Aware Intervention for On-Policy Distillation of Multi-Turn Agents." https://omanscience.com/en/articles/uopd-uncertainty-aware-intervention-for-on-policy-distillation-of-multi-turn-agents.
Chicago (author–date)
Zhang, Wenbo, Pengcheng Xu, Weizhi Du, Jing Zhang, and Hengrui Cai. 2026. "UOPD: Uncertainty-Aware Intervention for On-Policy Distillation of Multi-Turn Agents." https://omanscience.com/en/articles/uopd-uncertainty-aware-intervention-for-on-policy-distillation-of-multi-turn-agents.
Harvard
Zhang, W., Xu, P., Du, W., Zhang, J. and Cai, H. (2026) 'UOPD: Uncertainty-Aware Intervention for On-Policy Distillation of Multi-Turn Agents', Available at: https://omanscience.com/en/articles/uopd-uncertainty-aware-intervention-for-on-policy-distillation-of-multi-turn-agents.
Vancouver
Zhang W, Xu P, Du W, Zhang J, Cai H. UOPD: Uncertainty-Aware Intervention for On-Policy Distillation of Multi-Turn Agents. https://omanscience.com/en/articles/uopd-uncertainty-aware-intervention-for-on-policy-distillation-of-multi-turn-agents
IEEE
W. Zhang, P. Xu, W. Du, J. Zhang, and H. Cai, "UOPD: Uncertainty-Aware Intervention for On-Policy Distillation of Multi-Turn Agents," https://omanscience.com/en/articles/uopd-uncertainty-aware-intervention-for-on-policy-distillation-of-multi-turn-agents.