[
    {
        "id": "osp-22589",
        "type": "article-journal",
        "title": "Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering",
        "author": [
            {
                "family": "Jung",
                "given": "Jaewoo"
            },
            {
                "family": "Yu",
                "given": "Hyeonseo"
            },
            {
                "family": "An",
                "given": "Honggyu"
            },
            {
                "family": "Han",
                "given": "Jisang"
            },
            {
                "family": "Kim",
                "given": "Mungyeom"
            },
            {
                "family": "Jeon",
                "given": "Minkyeong"
            },
            {
                "family": "Shin",
                "given": "Heeseong"
            },
            {
                "family": "Moon",
                "given": "Wonjun"
            },
            {
                "family": "Tombari",
                "given": "Federico"
            },
            {
                "family": "Barath",
                "given": "Daniel"
            },
            {
                "family": "Pollefeys",
                "given": "Marc"
            },
            {
                "family": "Kim",
                "given": "Seungryong"
            },
            {
                "family": "Hong",
                "given": "Sunghwan"
            }
        ],
        "URL": "https://omanscience.com/ar/articles/imagine3d-llm-teaching-mllms-to-imagine-3d-scenes-before-answering",
        "language": "en",
        "issued": {
            "date-parts": [
                [
                    2026
                ]
            ]
        },
        "abstract": "Reasoning about the 3D world from multi-view images remains a fundamental challenge for Multimodal Large Language Models (MLLMs). While modern MLLMs handle single-image inputs effectively, they struggle to integrate evidence across viewpoints into a coherent 3D understanding. A growing body of work attempts to close this gap by injecting 3D awareness into MLLMs, either by boosting fine-grained pixel-level cross-view correspondence or by fusing features from 3D geometry foundation models, yet a substantial gap to human reasoning persists. In this work, we revisit human spatial reasoning, which suggests that rather than relying on fine-grained geometry cues, humans roughly identify common objects across views, infer the relative geometry between viewpoints, and assemble a coarse 3D layout of the scene. Inspired by this process, we introduce Imagine3D-LLM, an MLLM that learns to assemble a similar compact 3D representation of the scene and conditions its answer on this representation. Concretely, we append a small set of learnable summary tokens after the image tokens, decode them into a compact 3D Gaussian Splatting representation supervised by a photometric reconstruction loss, and train jointly with the standard next-token prediction objective. Notably, although only the summary tokens receive direct reconstruction supervision, this objective also induces stronger cross-frame correspondence within the LLM's underlying image features, suggesting that learning to reconstruct propagates 3D-aware signals throughout the model. As a result, Imagine3D-LLM consistently outperforms prior approaches across multiple spatial reasoning and 3D understanding benchmarks, suggesting that imagining the scene can be more effective than being told its pixel-wise geometry."
    }
]