Authors

Mengchao Zhang

Publications 2

Preprint Open access

SUAVE: Unified Video-Action Models via Masked Diffusion

Vision-language-action models (VLAs) inherit strong semantic grounding from pretrained vision-language backbones but are typically optimized for predicting actions rather than future observations. They can see and act, but they do not imagine the future before acting. World action models (WAMs) built on video diffusion …

Co-authors