الباحثون

Manling Li

المنشورات 3

نسخة أولية وصول مفتوح

WOVEN: Weaving Visual World Modeling into Multimodal LLMs

Zheyu Fan, Yue Zhang, Mingkai Deng وآخرون · 2026

Multimodal large language models (MLLMs) struggle with spatial, embodied, physical, and temporal reasoning. We hypothesize that these failures reflect a shared deficit in visual transition reasoning, and test whether this capability can serve as a shared training primitive, one that different models can learn from diff …

نسخة أولية وصول مفتوح

Representation-guided in-context learning for medical image interpretation with multimodal large language models

Minda Zhao, Fangyu Hu, Yan Luo وآخرون · 2026

Medical image interpretation is central to diagnosis and care, yet adapting general-purpose multimodal large language models (MLLMs) often requires resource-intensive domain-specific fine-tuning. Here we introduce representation-guided in-context learning (RG-ICL), a training-free inference framework that retrieves que …

المؤلفون المشاركون