Abstract
World action models (WAMs) jointly predict actions (intent) and visual future (foresight). Standard training adds noise to recorded actions and video simultaneously, but such training paradigms introduce a mismatch: perturbed actions imply counterfactual future visual, while the noised video remains tied to the GT recording. In low-noise regime, the scene geometry and even the dynamic behavior remain clearly visible from the noisy future frames despite the added noise. We present CtrlWAM, which executes perturbed actions in a simulator and pairs them with their noised visual consequences for joint WAM learning. To accommodate the different denoising requirements of video and actions, we introduce warped video--action noise schedules that aim to keep visual layout responsive as action predictions evolve. We further extend the action interface from ego-only control to a variable number of agent streams, allowing a unified model to represent predicted or commanded futures for multiple agents. Driving experiments show more accurate action forecasts, closer agreement between generated video and actions, and better following of supplied commands; robotics experiments show stronger motion fidelity and controllability. Matched controls support the benefit of off-path renders for command following and manipulation fidelity. Together, these findings contribute to a more controllable world action model. Project page: https://ctrl-wam.github.io/
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Peng, C., Ding, W., Tian, R., Zhou, Z., Packer, J., Igl, M., Karkus, P., Wang, Y., Tomizuka, M., Ivanovic, B., Pavone, M., & Chen, Y. (2026). CtrlWAM: Controllable World Action Models with Aligned Intent and Foresight. https://omanscience.com/en/articles/ctrlwam-controllable-world-action-models-with-aligned-intent-and-foresight
MLA 9
Peng, Chensheng, et al. "CtrlWAM: Controllable World Action Models with Aligned Intent and Foresight." https://omanscience.com/en/articles/ctrlwam-controllable-world-action-models-with-aligned-intent-and-foresight.
Chicago (author–date)
Peng, Chensheng, Wenhao Ding, Ran Tian, Zewei Zhou, Jef Packer, Maximilian Igl, Peter Karkus, Yan Wang, Masayoshi Tomizuka, Boris Ivanovic, Marco Pavone, and Yuxiao Chen. 2026. "CtrlWAM: Controllable World Action Models with Aligned Intent and Foresight." https://omanscience.com/en/articles/ctrlwam-controllable-world-action-models-with-aligned-intent-and-foresight.
Harvard
Peng, C., Ding, W., Tian, R., Zhou, Z., Packer, J., Igl, M., Karkus, P., Wang, Y., Tomizuka, M., Ivanovic, B., Pavone, M. and Chen, Y. (2026) 'CtrlWAM: Controllable World Action Models with Aligned Intent and Foresight', Available at: https://omanscience.com/en/articles/ctrlwam-controllable-world-action-models-with-aligned-intent-and-foresight.
Vancouver
Peng C, Ding W, Tian R, Zhou Z, Packer J, Igl M, et al. CtrlWAM: Controllable World Action Models with Aligned Intent and Foresight. https://omanscience.com/en/articles/ctrlwam-controllable-world-action-models-with-aligned-intent-and-foresight
IEEE
C. Peng, W. Ding, R. Tian, Z. Zhou, J. Packer, M. Igl, P. Karkus, Y. Wang, M. Tomizuka, B. Ivanovic, M. Pavone, and Y. Chen, "CtrlWAM: Controllable World Action Models with Aligned Intent and Foresight," https://omanscience.com/en/articles/ctrlwam-controllable-world-action-models-with-aligned-intent-and-foresight.