الباحثون

Sibo Song

المنشورات 1

نسخة أولية وصول مفتوح

VisionWeave: Weaving Elastic Visual Representations as a Native Capability of MLLMs

Yuan Feng, Qize Yang, Ruizhe Chen وآخرون · 2026

Multimodal large language models have become the dominant paradigm for visual understanding, but incur substantial costs by encoding inputs into dense, fixed-size patch tokens. However, visual information is unevenly distributed: some regions require fine-grained detail, while others admit compact representations. Down …

المؤلفون المشاركون