Abstract
World-action models (WAMs) jointly predict how a scene will evolve and how an agent should act, however joint generation alone does not necessarily impose a shared geometric constraint on these predictions. We present PhysWAM, a unified world-action model for autonomous driving that co-denoises multiview video, metric depth, and ego motion within a single flow-matching transformer. To ground world and action generation in measured scene geometry, we introduce Coupled Point Projection (CPP) that unprojects the generated depth into 3D points, transforms them using the generated $\mathrm{SE}(3)$ ego motion, and minimizes their distance to LiDAR points transformed using the recorded ego motion. This geometric constraint promotes physical consistency with the measured scene by jointly supervising generated depth and motion alongside their standard flow-matching objectives. At inference, trajectory selection relies only on a simple label-free consensus rule, with no learned scorer or simulator feedback. We evaluate PhysWAM across NAVSIM v1 and v2 planning, zero-shot closed-loop transfer, and future video and metric-depth prediction. Despite PhysWAM's simple selection procedure, it achieves strong planning performance and transfers zero-shot to unseen driving environments. It also generates accurate metric depth and temporally coherent video, with CPP improving both planning and depth prediction. Together, these results demonstrate that the geometric relationship between scene depth and ego motion provides a direct way to couple world and action generation within a simple unified model.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Parikh, D., Yu, F., Gao, Q., Yang, J., Ye, J., Bhatt, M., Vu, T., Ochoa, C., McAllister, R., Vasiljevic, I., Kannan, R., Prasanna, V., Guizilini, V., & Wang, Y. (2026). PhysWAM: Physically Consistent World Action Model for Autonomous Driving. https://omanscience.com/en/articles/physwam-physically-consistent-world-action-model-for-autonomous-driving
MLA 9
Parikh, Dhruv, et al. "PhysWAM: Physically Consistent World Action Model for Autonomous Driving." https://omanscience.com/en/articles/physwam-physically-consistent-world-action-model-for-autonomous-driving.
Chicago (author–date)
Parikh, Dhruv, Fengcheng Yu, Quankai Gao, Jiawei Yang, Junjie Ye, Maulik Bhatt, Thang Vu, Charles Ochoa, Rowan McAllister, Igor Vasiljevic, Rajgopal Kannan, Viktor Prasanna, Vitor Guizilini, and Yue Wang. 2026. "PhysWAM: Physically Consistent World Action Model for Autonomous Driving." https://omanscience.com/en/articles/physwam-physically-consistent-world-action-model-for-autonomous-driving.
Harvard
Parikh, D., Yu, F., Gao, Q., Yang, J., Ye, J., Bhatt, M., Vu, T., Ochoa, C., McAllister, R., Vasiljevic, I., Kannan, R., Prasanna, V., Guizilini, V. and Wang, Y. (2026) 'PhysWAM: Physically Consistent World Action Model for Autonomous Driving', Available at: https://omanscience.com/en/articles/physwam-physically-consistent-world-action-model-for-autonomous-driving.
Vancouver
Parikh D, Yu F, Gao Q, Yang J, Ye J, Bhatt M, et al. PhysWAM: Physically Consistent World Action Model for Autonomous Driving. https://omanscience.com/en/articles/physwam-physically-consistent-world-action-model-for-autonomous-driving
IEEE
D. Parikh, F. Yu, Q. Gao, J. Yang, J. Ye, M. Bhatt, T. Vu, C. Ochoa, R. McAllister, I. Vasiljevic, R. Kannan, V. Prasanna, V. Guizilini, and Y. Wang, "PhysWAM: Physically Consistent World Action Model for Autonomous Driving," https://omanscience.com/en/articles/physwam-physically-consistent-world-action-model-for-autonomous-driving.