Abstract

General-purpose vision-language models (VLMs) bring broad knowledge and spatial reasoning to robot manipulation, yet existing systems either use them indirectly, to predict constraints or write programs, or give them a view of the scene rather than a world in which to act. We present World Action Agent (WAA), a multi-agent harness through which VLMs pilot robots with basic tools, making every decision within a visual action workspace. The workspace has three properties. Contact views, selected automatically from the scene geometry, present the scene around the current interaction. Action rehearsal turns each action into an editable proposal that the agent, alone or through an Imagination Agent, previews and revises against planning feedback before execution. In-view correction closes the loop between observation, rehearsal, and low-level execution, letting the agent remove residual offsets in the view where it observes them. Through the same workspace, WAA acquires embodied procedural knowledge in two ways: it evolves multimodal skills from expert videos and human teaching under evidence-based review and consults them through a Skill Agent, and its interaction traces train smaller VLMs to pilot the same harness. On LIBERO-Pro, WAA with skills evolved only from LIBERO-90 reaches a state-of-the-art 75.6% average success, outperforming end-to-end VLAs, code-as-policy agents, and a visual-harness baseline with the same backbone; the same skills remain effective on robosuite without further learning. Fine-tuning Qwen3.5-9B on harness traces raises its out-of-domain success from 1.7% to 43.3%.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Zhang, Y., Huang, H., Chang, Y., Su, J., Zhou, B., Xu, Y., Chen, W., Zhou, T., Wang, C., Zhang, T., Wei, Y., Li, W., Deng, S., Li, Y., Chen, Y. C., & Li, Z. (2026). World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal. https://omanscience.com/en/articles/world-action-agent-harnessing-vlms-for-robot-manipulation-via-world-action-rehearsal

MLA 9

Zhang, Yehang, et al. "World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal." https://omanscience.com/en/articles/world-action-agent-harnessing-vlms-for-robot-manipulation-via-world-action-rehearsal.

Chicago (author–date)

Zhang, Yehang, Haojian Huang, Yifan Chang, Jianchong Su, Bohan Zhou, Yingjie Xu, Wosong Chen, Tianhao Zhou, Chenxu Wang, Tianyi Zhang, Yangkai Wei, Wenqian Li, Shiyuan Deng, Yinchuan Li, Ying-Cong Chen, and Zexi Li. 2026. "World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal." https://omanscience.com/en/articles/world-action-agent-harnessing-vlms-for-robot-manipulation-via-world-action-rehearsal.

Harvard

Zhang, Y., Huang, H., Chang, Y., Su, J., Zhou, B., Xu, Y., Chen, W., Zhou, T., Wang, C., Zhang, T., Wei, Y., Li, W., Deng, S., Li, Y., Chen, Y. C. and Li, Z. (2026) 'World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal', Available at: https://omanscience.com/en/articles/world-action-agent-harnessing-vlms-for-robot-manipulation-via-world-action-rehearsal.

Vancouver

Zhang Y, Huang H, Chang Y, Su J, Zhou B, Xu Y, et al. World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal. https://omanscience.com/en/articles/world-action-agent-harnessing-vlms-for-robot-manipulation-via-world-action-rehearsal

IEEE

Y. Zhang, H. Huang, Y. Chang, J. Su, B. Zhou, Y. Xu, W. Chen, T. Zhou, C. Wang, T. Zhang, Y. Wei, W. Li, S. Deng, Y. Li, Y. C. Chen, and Z. Li, "World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal," https://omanscience.com/en/articles/world-action-agent-harnessing-vlms-for-robot-manipulation-via-world-action-rehearsal.