Abstract

Predicting future spatial states supports collision avoidance and timely decision-making in dynamic environments. However, existing vision-language models (VLMs) and benchmarks for spatial reasoning primarily focus on observed scenes, leaving predictive spatial reasoning beyond the observed interval underexplored. To this end, we introduce SpatialMind, a metric-scale VLM for spatial reasoning and future prediction. Its metric depth adapter anchors spatial reasoning to real-world scale, while its progressive state chain establishes current spatial states and observed dynamics as the foundation for future prediction. Given a video prefix, SpatialMind predicts distances, motion directions, and spatial relations in both observed and unseen future frames. For training and evaluation, we build a scalable data engine that grounds entity descriptions in metric geometry to generate question-answer pairs and state supervision. Using this engine, we construct the SpatialMind-30K dataset and the SpatialMind-2K benchmark, both covering driving and everyday egocentric scenes. The benchmark spans eight tasks across three levels: current-state understanding, observed-dynamics understanding, and future prediction. Experiments show that SpatialMind substantially outperforms both general and spatially specialized models on our benchmark while achieving competitive zero-shot performance on VSI-Bench, OSI-Bench, and VLM4D.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Wang, F., Wang, X., Li, Z., He, W., Yan, Y., & Ren, L. (2026). From Sight to Foresight: Predictive Spatial Reasoning in Vision-Language Models. https://omanscience.com/en/articles/from-sight-to-foresight-predictive-spatial-reasoning-in-vision-language-models

MLA 9

Wang, Feiran, et al. "From Sight to Foresight: Predictive Spatial Reasoning in Vision-Language Models." https://omanscience.com/en/articles/from-sight-to-foresight-predictive-spatial-reasoning-in-vision-language-models.

Chicago (author–date)

Wang, Feiran, Xiaoqi Wang, Ziwei Li, Wenbin He, Yan Yan, and Liu Ren. 2026. "From Sight to Foresight: Predictive Spatial Reasoning in Vision-Language Models." https://omanscience.com/en/articles/from-sight-to-foresight-predictive-spatial-reasoning-in-vision-language-models.

Harvard

Wang, F., Wang, X., Li, Z., He, W., Yan, Y. and Ren, L. (2026) 'From Sight to Foresight: Predictive Spatial Reasoning in Vision-Language Models', Available at: https://omanscience.com/en/articles/from-sight-to-foresight-predictive-spatial-reasoning-in-vision-language-models.

Vancouver

Wang F, Wang X, Li Z, He W, Yan Y, Ren L. From Sight to Foresight: Predictive Spatial Reasoning in Vision-Language Models. https://omanscience.com/en/articles/from-sight-to-foresight-predictive-spatial-reasoning-in-vision-language-models

IEEE

F. Wang, X. Wang, Z. Li, W. He, Y. Yan, and L. Ren, "From Sight to Foresight: Predictive Spatial Reasoning in Vision-Language Models," https://omanscience.com/en/articles/from-sight-to-foresight-predictive-spatial-reasoning-in-vision-language-models.