Abstract

World action models (WAMs) jointly predict actions (intent) and visual future (foresight). Standard training adds noise to recorded actions and video simultaneously, but such training paradigms introduce a mismatch: perturbed actions imply counterfactual future visual, while the noised video remains tied to the GT recording. In low-noise regime, the scene geometry and even the dynamic behavior remain clearly visible from the noisy future frames despite the added noise. We present CtrlWAM, which executes perturbed actions in a simulator and pairs them with their noised visual consequences for joint WAM learning. To accommodate the different denoising requirements of video and actions, we introduce warped video--action noise schedules that aim to keep visual layout responsive as action predictions evolve. We further extend the action interface from ego-only control to a variable number of agent streams, allowing a unified model to represent predicted or commanded futures for multiple agents. Driving experiments show more accurate action forecasts, closer agreement between generated video and actions, and better following of supplied commands; robotics experiments show stronger motion fidelity and controllability. Matched controls support the benefit of off-path renders for command following and manipulation fidelity. Together, these findings contribute to a more controllable world action model. Project page: https://ctrl-wam.github.io/

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Peng, C., Ding, W., Tian, R., Zhou, Z., Packer, J., Igl, M., Karkus, P., Wang, Y., Tomizuka, M., Ivanovic, B., Pavone, M., & Chen, Y. (2026). CtrlWAM: Controllable World Action Models with Aligned Intent and Foresight. https://omanscience.com/en/articles/ctrlwam-controllable-world-action-models-with-aligned-intent-and-foresight

MLA 9

Peng, Chensheng, et al. "CtrlWAM: Controllable World Action Models with Aligned Intent and Foresight." https://omanscience.com/en/articles/ctrlwam-controllable-world-action-models-with-aligned-intent-and-foresight.

Chicago (author–date)

Peng, Chensheng, Wenhao Ding, Ran Tian, Zewei Zhou, Jef Packer, Maximilian Igl, Peter Karkus, Yan Wang, Masayoshi Tomizuka, Boris Ivanovic, Marco Pavone, and Yuxiao Chen. 2026. "CtrlWAM: Controllable World Action Models with Aligned Intent and Foresight." https://omanscience.com/en/articles/ctrlwam-controllable-world-action-models-with-aligned-intent-and-foresight.

Harvard

Peng, C., Ding, W., Tian, R., Zhou, Z., Packer, J., Igl, M., Karkus, P., Wang, Y., Tomizuka, M., Ivanovic, B., Pavone, M. and Chen, Y. (2026) 'CtrlWAM: Controllable World Action Models with Aligned Intent and Foresight', Available at: https://omanscience.com/en/articles/ctrlwam-controllable-world-action-models-with-aligned-intent-and-foresight.

Vancouver

Peng C, Ding W, Tian R, Zhou Z, Packer J, Igl M, et al. CtrlWAM: Controllable World Action Models with Aligned Intent and Foresight. https://omanscience.com/en/articles/ctrlwam-controllable-world-action-models-with-aligned-intent-and-foresight

IEEE

C. Peng, W. Ding, R. Tian, Z. Zhou, J. Packer, M. Igl, P. Karkus, Y. Wang, M. Tomizuka, B. Ivanovic, M. Pavone, and Y. Chen, "CtrlWAM: Controllable World Action Models with Aligned Intent and Foresight," https://omanscience.com/en/articles/ctrlwam-controllable-world-action-models-with-aligned-intent-and-foresight.