[
    {
        "id": "osp-24679",
        "type": "article-journal",
        "title": "Seeing and Solving Are Not Enough for Vision-Language Models",
        "author": [
            {
                "family": "Wang",
                "given": "Ziheng"
            },
            {
                "family": "Xie",
                "given": "Mingxuan"
            },
            {
                "family": "Liu",
                "given": "Yilin"
            },
            {
                "family": "Wu",
                "given": "Dayan"
            },
            {
                "family": "Li",
                "given": "Yang"
            },
            {
                "family": "Dai",
                "given": "Pengwen"
            }
        ],
        "URL": "https://omanscience.com/ar/articles/seeing-and-solving-are-not-enough-for-vision-language-models",
        "language": "en",
        "issued": {
            "date-parts": [
                [
                    2026
                ]
            ]
        },
        "abstract": "Vision-language models (VLMs) answer visual questions by combining visual information extraction with downstream problem solving. We investigate a fundamental question: Does an incorrect answer necessarily reflect a failure in visual extraction or problem solving? A model may succeed at both abilities when tested separately yet still fail on the original multimodal question, a distinction that overall answer accuracy cannot reveal. To study this, we perform a question-level empirical analysis across multiple VLMs and visual domains. We define an exactly scorable task state (i.e., the visual information sufficient to solve a question) and use it to test whether the same model can extract the required state, solve the question from the ground-truth state, and answer the original multimodal question. We find that composition failures, where extraction and solving both succeed but direct answering fails, account for 17.7% to 75.6% of direct-answering errors across multiple VLMs and datasets. To address this failure mode, we introduce a simple yet effective method, termed State Realization Tuning (SRT). SRT fine-tunes LoRA adapters attached to the language-model layers while keeping the pretrained VLM weights frozen. It trains the model to output the ground-truth task state before the final answer in a single autoregressive response. SRT improves over standard supervised fine-tuning by 1.7 to 14.1 percentage points and repairs 92.5% to 98.1% of diagnosed composition failures. A single LoRA adapter trained with SRT also improves performance across substantially different task-state structures. Our work shows that having both visual extraction and problem-solving capabilities does not guarantee correct multimodal answering. Requiring the model to first output the visual information needed to solve the question can help bridge this gap."
    }
]