نسخة أولية وصول مفتوح
Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering
Reasoning about the 3D world from multi-view images remains a fundamental challenge for Multimodal Large Language Models (MLLMs). While modern MLLMs handle single-image inputs effectively, they struggle to integrate evidence across viewpoints into a coherent 3D understanding. A growing body of work attempts to close th …