الملخص
Omnidirectional or 360 cameras provide embodied AI agents with a holistic, wide field-of-view (FoV) view of their surroundings, motivating the use of Multi-modal Large Language Models (MLLMs) for omnidirectional spatial reasoning. However, most MLLMs are trained on conventional 2D perspective images and struggle with the severe distortions and wrap-around discontinuities induced by spherical geometry. Enabling them to generalize to non-Euclidean 3D spaces without retraining therefore remains challenging. We propose SphMind, a training-free, plug-and-play framework that decouples semantic perception from geometric reasoning. Rather than requiring MLLMs to learn spherical geometry internally, SphMind preserves their semantic capabilities while handling geometry externally. We introduce a Spherical Harmonics-based Spatial Graph (SHSG) that models spatial relationships through equivariant transformations on the sphere, together with Inference-Time Geometric Grounding (IGG), a model-agnostic closed-loop optimization process that aligns MLLM representations with spherical geometric constraints during inference. Experiments on three benchmarks show that SphMind achieves over 21.4% average improvement in directional reasoning on MP3D and Stanford2D-3D, outperforms prompt-engineering baselines by 8.7% on the real-world ODI-Bench, and improves rotational invariance by 5.9% under panorama rotations, without additional training or dataset-specific tuning. In-the-wild evaluations further show that SphMind resolves directional reasoning queries that baseline vision-language models fail to answer correctly.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Damodaran, S., Debnath, S., Tan, C., & Wang, L. (2026). SphMind: Towards Robust, Training-Free VLM-based Spatial Reasoning with a 360 Camera. https://omanscience.com/ar/articles/sphmind-towards-robust-training-free-vlm-based-spatial-reasoning-with-a-360-camera
MLA 9
Damodaran, Shriram, et al. "SphMind: Towards Robust, Training-Free VLM-based Spatial Reasoning with a 360 Camera." https://omanscience.com/ar/articles/sphmind-towards-robust-training-free-vlm-based-spatial-reasoning-with-a-360-camera.
شيكاغو (المؤلف–التاريخ)
Damodaran, Shriram, Soumyaratna Debnath, Cheston Tan, and Lin Wang. 2026. "SphMind: Towards Robust, Training-Free VLM-based Spatial Reasoning with a 360 Camera." https://omanscience.com/ar/articles/sphmind-towards-robust-training-free-vlm-based-spatial-reasoning-with-a-360-camera.
هارفارد
Damodaran, S., Debnath, S., Tan, C. and Wang, L. (2026) 'SphMind: Towards Robust, Training-Free VLM-based Spatial Reasoning with a 360 Camera', Available at: https://omanscience.com/ar/articles/sphmind-towards-robust-training-free-vlm-based-spatial-reasoning-with-a-360-camera.
فانكوفر
Damodaran S, Debnath S, Tan C, Wang L. SphMind: Towards Robust, Training-Free VLM-based Spatial Reasoning with a 360 Camera. https://omanscience.com/ar/articles/sphmind-towards-robust-training-free-vlm-based-spatial-reasoning-with-a-360-camera
IEEE
S. Damodaran, S. Debnath, C. Tan, and L. Wang, "SphMind: Towards Robust, Training-Free VLM-based Spatial Reasoning with a 360 Camera," https://omanscience.com/ar/articles/sphmind-towards-robust-training-free-vlm-based-spatial-reasoning-with-a-360-camera.