الباحثون

Weilun Feng

المنشورات 2

نسخة أولية وصول مفتوح

SceneScaffold: Active Scene-State Construction for Unified 3D Scene Understanding

Xiangqi Li, Libo Huang, Jiarui Zhao وآخرون · 2026

Recent 3D large multimodal models (3D-LMMs) rely on a visual bottleneck to compress complex 3D scene evidence into a limited number of visual tokens compatible with large language models (LLMs). Current visual bottlenecks, however, often passively compress heterogeneous 3D evidence into a homogeneous object-centric tok …

المؤلفون المشاركون