Abstract
Predicting future spatial states supports collision avoidance and timely decision-making in dynamic environments. However, existing vision-language models (VLMs) and benchmarks for spatial reasoning primarily focus on observed scenes, leaving predictive spatial reasoning beyond the observed interval underexplored. To this end, we introduce SpatialMind, a metric-scale VLM for spatial reasoning and future prediction. Its metric depth adapter anchors spatial reasoning to real-world scale, while its progressive state chain establishes current spatial states and observed dynamics as the foundation for future prediction. Given a video prefix, SpatialMind predicts distances, motion directions, and spatial relations in both observed and unseen future frames. For training and evaluation, we build a scalable data engine that grounds entity descriptions in metric geometry to generate question-answer pairs and state supervision. Using this engine, we construct the SpatialMind-30K dataset and the SpatialMind-2K benchmark, both covering driving and everyday egocentric scenes. The benchmark spans eight tasks across three levels: current-state understanding, observed-dynamics understanding, and future prediction. Experiments show that SpatialMind substantially outperforms both general and spatially specialized models on our benchmark while achieving competitive zero-shot performance on VSI-Bench, OSI-Bench, and VLM4D.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Wang, F., Wang, X., Li, Z., He, W., Yan, Y., & Ren, L. (2026). From Sight to Foresight: Predictive Spatial Reasoning in Vision-Language Models. https://omanscience.com/en/articles/from-sight-to-foresight-predictive-spatial-reasoning-in-vision-language-models
MLA 9
Wang, Feiran, et al. "From Sight to Foresight: Predictive Spatial Reasoning in Vision-Language Models." https://omanscience.com/en/articles/from-sight-to-foresight-predictive-spatial-reasoning-in-vision-language-models.
Chicago (author–date)
Wang, Feiran, Xiaoqi Wang, Ziwei Li, Wenbin He, Yan Yan, and Liu Ren. 2026. "From Sight to Foresight: Predictive Spatial Reasoning in Vision-Language Models." https://omanscience.com/en/articles/from-sight-to-foresight-predictive-spatial-reasoning-in-vision-language-models.
Harvard
Wang, F., Wang, X., Li, Z., He, W., Yan, Y. and Ren, L. (2026) 'From Sight to Foresight: Predictive Spatial Reasoning in Vision-Language Models', Available at: https://omanscience.com/en/articles/from-sight-to-foresight-predictive-spatial-reasoning-in-vision-language-models.
Vancouver
Wang F, Wang X, Li Z, He W, Yan Y, Ren L. From Sight to Foresight: Predictive Spatial Reasoning in Vision-Language Models. https://omanscience.com/en/articles/from-sight-to-foresight-predictive-spatial-reasoning-in-vision-language-models
IEEE
F. Wang, X. Wang, Z. Li, W. He, Y. Yan, and L. Ren, "From Sight to Foresight: Predictive Spatial Reasoning in Vision-Language Models," https://omanscience.com/en/articles/from-sight-to-foresight-predictive-spatial-reasoning-in-vision-language-models.