الملخص

تمت ترجمة أجزاء من هذه الصفحة آلياً وقد تحتوي على أخطاء.

A multimodal large language model that correctly reports a spatial fact does not necessarily use that fact in subsequent reasoning. To study this distinction, we introduce \textsc{SpaceConflict}, a benchmark of 23{,}196 inputs for the construction and use of spatial state. Under a unified Supported/Contradictory/Unknown judgment interface, it covers local fact binding (L1), relational composition (L2), cross-observation consistency (L3), and state judgment under transformation (L4). Posing a direct-state query, a full-transformation query, and an explicit-initial-state query on the same world reveals an availability--utilization gap: models recover the initial state from visual evidence yet fail when that state must drive a transformation. For Qwen3.5-9B, 50 of 100 sequences with a correctly recovered initial state fail the full transformation, and supplying the state explicitly repairs all 50; the gap narrows with scale but does not close. We therefore propose Operational State Supervision (OSS), which supervises task-relevant spatial states and their transformation trajectories and aligns shared facts across contexts. OSS improves paired accuracy on matched judgments most on L3 and L4, the levels that depend on organizing and using state. Evaluating multimodal spatial reasoning thus requires asking not only whether a model can see and state a spatial fact, but whether that fact becomes a usable state in subsequent computation.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Zhang, J., & Lu, G. (2026). الرؤية والقول دون الاستخدام: من الحقائق المكانية القابلة للتصريح إلى الحالات القابلة للاستخدام في نماذج اللغة الكبيرة متعددة الوسائط. https://omanscience.com/ar/articles/seeing-saying-but-not-using-from-reportable-spatial-facts-to-usable-states-in-multimodal-large-language-models

MLA 9

Zhang, Jinchang, and Guoyu Lu. "الرؤية والقول دون الاستخدام: من الحقائق المكانية القابلة للتصريح إلى الحالات القابلة للاستخدام في نماذج اللغة الكبيرة متعددة الوسائط." https://omanscience.com/ar/articles/seeing-saying-but-not-using-from-reportable-spatial-facts-to-usable-states-in-multimodal-large-language-models.

شيكاغو (المؤلف–التاريخ)

Zhang, Jinchang, and Guoyu Lu. 2026. "الرؤية والقول دون الاستخدام: من الحقائق المكانية القابلة للتصريح إلى الحالات القابلة للاستخدام في نماذج اللغة الكبيرة متعددة الوسائط." https://omanscience.com/ar/articles/seeing-saying-but-not-using-from-reportable-spatial-facts-to-usable-states-in-multimodal-large-language-models.

هارفارد

Zhang, J. and Lu, G. (2026) 'الرؤية والقول دون الاستخدام: من الحقائق المكانية القابلة للتصريح إلى الحالات القابلة للاستخدام في نماذج اللغة الكبيرة متعددة الوسائط', Available at: https://omanscience.com/ar/articles/seeing-saying-but-not-using-from-reportable-spatial-facts-to-usable-states-in-multimodal-large-language-models.

فانكوفر

Zhang J, Lu G. الرؤية والقول دون الاستخدام: من الحقائق المكانية القابلة للتصريح إلى الحالات القابلة للاستخدام في نماذج اللغة الكبيرة متعددة الوسائط. https://omanscience.com/ar/articles/seeing-saying-but-not-using-from-reportable-spatial-facts-to-usable-states-in-multimodal-large-language-models

IEEE

J. Zhang, and G. Lu, "الرؤية والقول دون الاستخدام: من الحقائق المكانية القابلة للتصريح إلى الحالات القابلة للاستخدام في نماذج اللغة الكبيرة متعددة الوسائط," https://omanscience.com/ar/articles/seeing-saying-but-not-using-from-reportable-spatial-facts-to-usable-states-in-multimodal-large-language-models.