الملخص
Existing approaches to multi-view spatial reasoning operate largely on sparse input views. Vision-language models (VLMs) are thus restricted to understand a scene and infer spatial relations within these fixed views, leading to fragile cross-view alignment and geometry-to-language bottleneck. To address these issues, we formulate a novel Seek-and-View reasoning approach to find implicit cross-view spatial evidence by locating a question-relevant view to support the spatial reasoning. To realize this approach, we propose Vantage, a training-free model-agnostic reasoning framework that pairs a VLM with a 3D foundation model: a viewpoint-grounded reasoning stage for question analysis and view planning, followed by a geometry-grounded evidence augmentation stage to effectively synthesize and incorporate visual evidence into the final reasoning. Comprehensive experiments on six VLMs demonstrate consistent improvements on five benchmarks without fine-tuning. Overall, by revealing spatial evidence through view-grounded reasoning, Vantage can largely reduce reliance on language-based cross-view alignment and improve multi-view spatial understanding. Our code is available at https://github.com/q1xiangchen/Vantage.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Chen, Q., Zhang, C., Ke, F., Fu, C. W., Cai, J., & Ye, J. (2026). Seek-and-View Reasoning for Multi-View Spatial Understanding. https://omanscience.com/ar/articles/seek-and-view-reasoning-for-multi-view-spatial-understanding
MLA 9
Chen, Qixiang, et al. "Seek-and-View Reasoning for Multi-View Spatial Understanding." https://omanscience.com/ar/articles/seek-and-view-reasoning-for-multi-view-spatial-understanding.
شيكاغو (المؤلف–التاريخ)
Chen, Qixiang, Cheng Zhang, Fucai Ke, Chi-Wing Fu, Jianfei Cai, and Jingwen Ye. 2026. "Seek-and-View Reasoning for Multi-View Spatial Understanding." https://omanscience.com/ar/articles/seek-and-view-reasoning-for-multi-view-spatial-understanding.
هارفارد
Chen, Q., Zhang, C., Ke, F., Fu, C. W., Cai, J. and Ye, J. (2026) 'Seek-and-View Reasoning for Multi-View Spatial Understanding', Available at: https://omanscience.com/ar/articles/seek-and-view-reasoning-for-multi-view-spatial-understanding.
فانكوفر
Chen Q, Zhang C, Ke F, Fu CW, Cai J, Ye J. Seek-and-View Reasoning for Multi-View Spatial Understanding. https://omanscience.com/ar/articles/seek-and-view-reasoning-for-multi-view-spatial-understanding
IEEE
Q. Chen, C. Zhang, F. Ke, C. W. Fu, J. Cai, and J. Ye, "Seek-and-View Reasoning for Multi-View Spatial Understanding," https://omanscience.com/ar/articles/seek-and-view-reasoning-for-multi-view-spatial-understanding.