Abstract
Online training enables computer-use agents (CUAs) to improve through interaction with executable environments. However, existing methods primarily rely on sparse outcome rewards, which provide no supervision for intermediate actions. On-policy self-distillation (OPSD) offers token-level learning signals through privileged rescoring, but directly applying it to CUA online training presents two challenges: fixed guidance may become misaligned with the student's current state, and guidance-induced probability shifts may conflict with step-level correctness. We introduce ComputerSD, an online self-distillation method for CUAs that converts real-time feedback from executed GUI transitions into guidance for policy learning. A fine-tuned GUI analyzer produces guidance and a step-level value score after each action; the guidance provides privileged context, while the score regulates the resulting OPSD signals. ComputerSD jointly optimizes token-level OPSD and trajectory-level GRPO in a fully asynchronous training framework. On OSWorld-Verified, ComputerSD outperforms outcome-only GRPO by 1.9 and 4.1 percentage points on the general-purpose Qwen3-VL-8B-Thinking and specialized EvoCUA-8B backbones, respectively. Evaluation in out-of-distribution settings further supports the generalizability of ComputerSD. These results demonstrate the effectiveness of learning from real-time feedback through online self-distillation for CUAs.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Du, Y., Chen, T., Lu, Z., Liu, Y., Chen, B., Jiang, T., Xu, W., & Shen, Y. (2026). ComputerSD: Online Self-Distillation from Real-Time Feedback for Computer-Use Agents. https://omanscience.com/en/articles/computersd-online-self-distillation-from-real-time-feedback-for-computer-use-agents
MLA 9
Du, Yong, et al. "ComputerSD: Online Self-Distillation from Real-Time Feedback for Computer-Use Agents." https://omanscience.com/en/articles/computersd-online-self-distillation-from-real-time-feedback-for-computer-use-agents.
Chicago (author–date)
Du, Yong, Tongbo Chen, Zhengxi Lu, Yizhou Liu, Bofan Chen, Tao Jiang, Wenhao Xu, and Yongliang Shen. 2026. "ComputerSD: Online Self-Distillation from Real-Time Feedback for Computer-Use Agents." https://omanscience.com/en/articles/computersd-online-self-distillation-from-real-time-feedback-for-computer-use-agents.
Harvard
Du, Y., Chen, T., Lu, Z., Liu, Y., Chen, B., Jiang, T., Xu, W. and Shen, Y. (2026) 'ComputerSD: Online Self-Distillation from Real-Time Feedback for Computer-Use Agents', Available at: https://omanscience.com/en/articles/computersd-online-self-distillation-from-real-time-feedback-for-computer-use-agents.
Vancouver
Du Y, Chen T, Lu Z, Liu Y, Chen B, Jiang T, et al. ComputerSD: Online Self-Distillation from Real-Time Feedback for Computer-Use Agents. https://omanscience.com/en/articles/computersd-online-self-distillation-from-real-time-feedback-for-computer-use-agents
IEEE
Y. Du, T. Chen, Z. Lu, Y. Liu, B. Chen, T. Jiang, W. Xu, and Y. Shen, "ComputerSD: Online Self-Distillation from Real-Time Feedback for Computer-Use Agents," https://omanscience.com/en/articles/computersd-online-self-distillation-from-real-time-feedback-for-computer-use-agents.