Abstract
World action models (WAMs) have emerged as a promising paradigm for robotic control by jointly predicting future visual dynamics and actions from an initial observation and instruction. However, existing WAMs struggle with long-horizon prediction, as generating dense video rollouts is highly inefficient. Some recent WAMs address this by predicting a single future frame without generating the full video, but this approach neglects how to progress toward the goal. We present ProWAM, a progressive world action model that jointly predicts actions and an ordered sequence of sparse visual sub-goals, providing explicit visual guidance to anchor action generation throughout task execution. This design scales naturally, as sub-goal prediction can be learned from large-scale action-free videos, allowing the video backbone to offload complex visual planning from the action policy. For efficient action generation, ProWAM executes a single video-backbone forward pass to cache sparse sub-goal features, eliminating iterative full-video generation and requiring only lightweight action denoising during replanning. Across extensive evaluations, ProWAM achieves superior out-of-distribution robustness. On simulation benchmarks, it sets new state-of-the-art results on LIBERO-Plus (85.8%) and randomized RoboTwin (75.7%), outperforming the strongest baseline with relative gains of up to +35.9%. On RoboCasa365, ProWAM achieves a 48.1% success rate and 18.2% on the challenging Composite-Unseen split, ranking 4th overall. Crucially, in zero-shot real-world experiments, ProWAM achieves 70.0% success, outperforming the strongest baseline by +15.0 (from 55.0% to 70.0%, a +27.3% relative gain) in novel scenes. These results demonstrate the value of progress-indexed visual foresight for closed-loop control. Our program is in https://sii-ferenas.github.io/ProWAM-page.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Zhang, F., An, Z., Frost, D., Wang, Y., Liu, P., Zhang, Y., Drozdzal, M., & Bar, A. (2026). World Action Modeling with Progressive Visual Planning. https://omanscience.com/en/articles/world-action-modeling-with-progressive-visual-planning
MLA 9
Zhang, Fei, et al. "World Action Modeling with Progressive Visual Planning." https://omanscience.com/en/articles/world-action-modeling-with-progressive-visual-planning.
Chicago (author–date)
Zhang, Fei, Zhaochong An, Duncan Frost, Yikai Wang, Pengfei Liu, Ya Zhang, Michal Drozdzal, and Amir Bar. 2026. "World Action Modeling with Progressive Visual Planning." https://omanscience.com/en/articles/world-action-modeling-with-progressive-visual-planning.
Harvard
Zhang, F., An, Z., Frost, D., Wang, Y., Liu, P., Zhang, Y., Drozdzal, M. and Bar, A. (2026) 'World Action Modeling with Progressive Visual Planning', Available at: https://omanscience.com/en/articles/world-action-modeling-with-progressive-visual-planning.
Vancouver
Zhang F, An Z, Frost D, Wang Y, Liu P, Zhang Y, et al. World Action Modeling with Progressive Visual Planning. https://omanscience.com/en/articles/world-action-modeling-with-progressive-visual-planning
IEEE
F. Zhang, Z. An, D. Frost, Y. Wang, P. Liu, Y. Zhang, M. Drozdzal, and A. Bar, "World Action Modeling with Progressive Visual Planning," https://omanscience.com/en/articles/world-action-modeling-with-progressive-visual-planning.