Authors

Canyu Chen

Publications 1

Preprint Open access

WOVEN: Weaving Visual World Modeling into Multimodal LLMs

Zheyu Fan, Yue Zhang, Mingkai Deng et al. · 2026

Multimodal large language models (MLLMs) struggle with spatial, embodied, physical, and temporal reasoning. We hypothesize that these failures reflect a shared deficit in visual transition reasoning, and test whether this capability can serve as a shared training primitive, one that different models can learn from diff …

Co-authors