الملخص

Vision-language models (VLMs) increasingly operate in embodied and spatially grounded settings, where accurate understanding of depth, viewpoint, and three-dimensional relations is essential. However, improving spatial reasoning typically relies on ground-truth answers, answer-derived rewards, or other forms of task-specific supervision. We introduce Spatial-OPSD, a label-free self-improvement framework that instead exploits spatial structure naturally available from perception and reconstruction tools. During training, a privileged teacher receives automatically obtainable spatial priors, such as depth, reconstructed 3D relations, and camera geometry, while the student observes only the original visual-language input. On trajectories sampled by the student itself, the teacher provides dense token-level supervision, allowing the student to internalize spatial knowledge without ground-truth answer labels or privileged information at inference time. To extend this supervision beyond a single round, we adopt a round-wise recursive training scheme: the teacher remains frozen within each round to provide a stable learning target, and the improved student initializes both teacher and student in the next round, where privileged spatial priors re-establish an informative teacher--student asymmetry. This enables repeated self-improvement while avoiding a rapidly moving teacher during optimization. Across four VLM families, a single round of Spatial-OPSD consistently improves the five-benchmark average, while three rounds further push a strong spatially specialized model to the open-source frontier, achieving the highest average among the open models and the best results on three of five spatial reasoning benchmarks. Our code is available at https://github.com/vermouth599/Spatial-OPSD.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Liu, Z., Chen, Z., Chen, K., Sun, M., An, X., Jing, H., & Huang, R. (2026). Spatial-OPSD: Self-Improving Spatial Reasoning via Label-Free Self-Distillation. https://omanscience.com/ar/articles/spatial-opsd-self-improving-spatial-reasoning-via-label-free-self-distillation

MLA 9

Liu, Zhenyu, et al. "Spatial-OPSD: Self-Improving Spatial Reasoning via Label-Free Self-Distillation." https://omanscience.com/ar/articles/spatial-opsd-self-improving-spatial-reasoning-via-label-free-self-distillation.

شيكاغو (المؤلف–التاريخ)

Liu, Zhenyu, Zhangquan Chen, Keyi Chen, Mingze Sun, Xiang An, Haodong Jing, and Ruqi Huang. 2026. "Spatial-OPSD: Self-Improving Spatial Reasoning via Label-Free Self-Distillation." https://omanscience.com/ar/articles/spatial-opsd-self-improving-spatial-reasoning-via-label-free-self-distillation.

هارفارد

Liu, Z., Chen, Z., Chen, K., Sun, M., An, X., Jing, H. and Huang, R. (2026) 'Spatial-OPSD: Self-Improving Spatial Reasoning via Label-Free Self-Distillation', Available at: https://omanscience.com/ar/articles/spatial-opsd-self-improving-spatial-reasoning-via-label-free-self-distillation.

فانكوفر

Liu Z, Chen Z, Chen K, Sun M, An X, Jing H, et al. Spatial-OPSD: Self-Improving Spatial Reasoning via Label-Free Self-Distillation. https://omanscience.com/ar/articles/spatial-opsd-self-improving-spatial-reasoning-via-label-free-self-distillation

IEEE

Z. Liu, Z. Chen, K. Chen, M. Sun, X. An, H. Jing, and R. Huang, "Spatial-OPSD: Self-Improving Spatial Reasoning via Label-Free Self-Distillation," https://omanscience.com/ar/articles/spatial-opsd-self-improving-spatial-reasoning-via-label-free-self-distillation.