Abstract
Large Vision-Language Models (LVLMs) have made remarkable progress across visual perception tasks, yet spatial reasoning remains a persistent weakness, especially for questions that require reasoning over visual space. Recent spatial-reasoning methods incorporate generated grounding, where models predict bounding boxes, masks, or other localization outputs for task-relevant objects as part of their reasoning trace. However, these approaches typically optimize final-answer correctness alone, allowing correct answers to be rewarded even when the model does not reason from confidently localized task-relevant objects. We introduce SpatialCORE (Spatially COnfident REasoning), a post-training framework that turns the model's own confidence in generated grounding into a learning signal for spatial reasoning. Its central idea is to reinforce grounding that is both accurate and confident, encouraging the model to reason from confidently localized task-relevant objects. SpatialCORE realizes this through a self-regulating spatial reward that weights each predicted bounding box's matching quality by its coordinate-token confidence. An answer gate further ties grounding optimization to final-answer correctness. SpatialCORE achieves state-of-the-art results among open-source and specialized spatial reasoning models across diverse benchmarks, and transfers effectively in zero-shot settings to unseen data distributions. The source code is available at https://github.com/rafiibnsultan/SpatialCORE.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Ibn Sultan, R., Zhou, X., Chowdhury, M. S. A., Li, C., Khanduri, P., Brocanelli, M., & Zhu, D. (2026). SpatialCORE: Confidence-Aware Grounded Spatial Reasoning in Large Vision--Language Models. https://omanscience.com/en/articles/spatialcore-confidence-aware-grounded-spatial-reasoning-in-large-vision-language-models
MLA 9
Ibn Sultan, Rafi, et al. "SpatialCORE: Confidence-Aware Grounded Spatial Reasoning in Large Vision--Language Models." https://omanscience.com/en/articles/spatialcore-confidence-aware-grounded-spatial-reasoning-in-large-vision-language-models.
Chicago (author–date)
Ibn Sultan, Rafi, Xiangyu Zhou, Md. Sajid Alam Chowdhury, Chengyin Li, Prashant Khanduri, Marco Brocanelli, and Dongxiao Zhu. 2026. "SpatialCORE: Confidence-Aware Grounded Spatial Reasoning in Large Vision--Language Models." https://omanscience.com/en/articles/spatialcore-confidence-aware-grounded-spatial-reasoning-in-large-vision-language-models.
Harvard
Ibn Sultan, R., Zhou, X., Chowdhury, M. S. A., Li, C., Khanduri, P., Brocanelli, M. and Zhu, D. (2026) 'SpatialCORE: Confidence-Aware Grounded Spatial Reasoning in Large Vision--Language Models', Available at: https://omanscience.com/en/articles/spatialcore-confidence-aware-grounded-spatial-reasoning-in-large-vision-language-models.
Vancouver
Ibn Sultan R, Zhou X, Chowdhury MSA, Li C, Khanduri P, Brocanelli M, et al. SpatialCORE: Confidence-Aware Grounded Spatial Reasoning in Large Vision--Language Models. https://omanscience.com/en/articles/spatialcore-confidence-aware-grounded-spatial-reasoning-in-large-vision-language-models
IEEE
R. Ibn Sultan, X. Zhou, M. S. A. Chowdhury, C. Li, P. Khanduri, M. Brocanelli, and D. Zhu, "SpatialCORE: Confidence-Aware Grounded Spatial Reasoning in Large Vision--Language Models," https://omanscience.com/en/articles/spatialcore-confidence-aware-grounded-spatial-reasoning-in-large-vision-language-models.