الملخص
Online training enables computer-use agents (CUAs) to improve through interaction with executable environments. However, existing methods primarily rely on sparse outcome rewards, which provide no supervision for intermediate actions. On-policy self-distillation (OPSD) offers token-level learning signals through privileged rescoring, but directly applying it to CUA online training presents two challenges: fixed guidance may become misaligned with the student's current state, and guidance-induced probability shifts may conflict with step-level correctness. We introduce ComputerSD, an online self-distillation method for CUAs that converts real-time feedback from executed GUI transitions into guidance for policy learning. A fine-tuned GUI analyzer produces guidance and a step-level value score after each action; the guidance provides privileged context, while the score regulates the resulting OPSD signals. ComputerSD jointly optimizes token-level OPSD and trajectory-level GRPO in a fully asynchronous training framework. On OSWorld-Verified, ComputerSD outperforms outcome-only GRPO by 1.9 and 4.1 percentage points on the general-purpose Qwen3-VL-8B-Thinking and specialized EvoCUA-8B backbones, respectively. Evaluation in out-of-distribution settings further supports the generalizability of ComputerSD. These results demonstrate the effectiveness of learning from real-time feedback through online self-distillation for CUAs.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Du, Y., Chen, T., Lu, Z., Liu, Y., Chen, B., Jiang, T., Xu, W., & Shen, Y. (2026). ComputerSD: Online Self-Distillation from Real-Time Feedback for Computer-Use Agents. https://omanscience.com/ar/articles/computersd-online-self-distillation-from-real-time-feedback-for-computer-use-agents
MLA 9
Du, Yong, et al. "ComputerSD: Online Self-Distillation from Real-Time Feedback for Computer-Use Agents." https://omanscience.com/ar/articles/computersd-online-self-distillation-from-real-time-feedback-for-computer-use-agents.
شيكاغو (المؤلف–التاريخ)
Du, Yong, Tongbo Chen, Zhengxi Lu, Yizhou Liu, Bofan Chen, Tao Jiang, Wenhao Xu, and Yongliang Shen. 2026. "ComputerSD: Online Self-Distillation from Real-Time Feedback for Computer-Use Agents." https://omanscience.com/ar/articles/computersd-online-self-distillation-from-real-time-feedback-for-computer-use-agents.
هارفارد
Du, Y., Chen, T., Lu, Z., Liu, Y., Chen, B., Jiang, T., Xu, W. and Shen, Y. (2026) 'ComputerSD: Online Self-Distillation from Real-Time Feedback for Computer-Use Agents', Available at: https://omanscience.com/ar/articles/computersd-online-self-distillation-from-real-time-feedback-for-computer-use-agents.
فانكوفر
Du Y, Chen T, Lu Z, Liu Y, Chen B, Jiang T, et al. ComputerSD: Online Self-Distillation from Real-Time Feedback for Computer-Use Agents. https://omanscience.com/ar/articles/computersd-online-self-distillation-from-real-time-feedback-for-computer-use-agents
IEEE
Y. Du, T. Chen, Z. Lu, Y. Liu, B. Chen, T. Jiang, W. Xu, and Y. Shen, "ComputerSD: Online Self-Distillation from Real-Time Feedback for Computer-Use Agents," https://omanscience.com/ar/articles/computersd-online-self-distillation-from-real-time-feedback-for-computer-use-agents.