Abstract

Improving robot policies on new tasks without collecting additional expert demonstrations remains a central challenge in robot learning. World action models (WAMs) use broad video priors to jointly predict future videos and actions, offering a potential source of supervision for adapting to new tasks. However, generated videos may fail to depict task completion, and even visually successful videos may be paired with inconsistent actions that lead to execution failure. We propose EVO-WAM, a framework that adapts WAMs to unseen tasks by learning from their own generated video-action trajectories, without executing candidate actions in an external environment. First, we augment WAM training with state prediction and anchored multi-frame context to enable complete autoregressive rollouts without external execution feedback. Second, we identify reliable training experience by selecting task-completing prefixes with a vision-language model and verifying their video-action consistency with an inverse dynamics model. Third, we iteratively train the WAM on verified prefixes and generate new rollouts with the updated model. On seven unseen RoboTwin 2.0 tasks, EVO-WAM increases average success rates from 26.9% to 68.0% for Cosmos3 and from 28.5% to 46.4% for DreamZero, reaching approximately $2.5\times$ and $1.6\times$ their initial success rates. On three unseen long-horizon composite tasks in the real world, it improves Cosmos3's average success rate from 20.0% to 76.7%, a gain of 56.7 percentage points. Project Page: https://evo-wam.github.io/.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Zhou, S., Wu, X., Li, W., Zheng, S., Zhang, J., Yu, S., Yang, Y., Liu, J., Sun, H., Yang, S., Jiang, L., Su, J., Huang, H., & Tian, Z. (2026). EVO-WAM: Evolving World Action Models through Video-Action Verification. https://omanscience.com/en/articles/evo-wam-evolving-world-action-models-through-video-action-verification

MLA 9

Zhou, Shiyang, et al. "EVO-WAM: Evolving World Action Models through Video-Action Verification." https://omanscience.com/en/articles/evo-wam-evolving-world-action-models-through-video-action-verification.

Chicago (author–date)

Zhou, Shiyang, Xionghao Wu, Wenbo Li, Shenghe Zheng, Jiyao Zhang, Songsong Yu, Yijun Yang, Jianhui Liu, Haoze Sun, Senqiao Yang, Li Jiang, Jingyong Su, Haoyang Huang, and Zhuotao Tian. 2026. "EVO-WAM: Evolving World Action Models through Video-Action Verification." https://omanscience.com/en/articles/evo-wam-evolving-world-action-models-through-video-action-verification.

Harvard

Zhou, S., Wu, X., Li, W., Zheng, S., Zhang, J., Yu, S., Yang, Y., Liu, J., Sun, H., Yang, S., Jiang, L., Su, J., Huang, H. and Tian, Z. (2026) 'EVO-WAM: Evolving World Action Models through Video-Action Verification', Available at: https://omanscience.com/en/articles/evo-wam-evolving-world-action-models-through-video-action-verification.

Vancouver

Zhou S, Wu X, Li W, Zheng S, Zhang J, Yu S, et al. EVO-WAM: Evolving World Action Models through Video-Action Verification. https://omanscience.com/en/articles/evo-wam-evolving-world-action-models-through-video-action-verification

IEEE

S. Zhou, X. Wu, W. Li, S. Zheng, J. Zhang, S. Yu, Y. Yang, J. Liu, H. Sun, S. Yang, L. Jiang, J. Su, H. Huang, and Z. Tian, "EVO-WAM: Evolving World Action Models through Video-Action Verification," https://omanscience.com/en/articles/evo-wam-evolving-world-action-models-through-video-action-verification.