الباحثون

Mingkai Deng

المنشورات 1

نسخة أولية وصول مفتوح

WOVEN: Weaving Visual World Modeling into Multimodal LLMs

Zheyu Fan, Yue Zhang, Mingkai Deng وآخرون · 2026

Multimodal large language models (MLLMs) struggle with spatial, embodied, physical, and temporal reasoning. We hypothesize that these failures reflect a shared deficit in visual transition reasoning, and test whether this capability can serve as a shared training primitive, one that different models can learn from diff …

المؤلفون المشاركون