الملخص
تمت ترجمة أجزاء من هذه الصفحة آلياً وقد تحتوي على أخطاء.
Recent works augment Vision-Language Models with geometry features from pretrained 3D models, expecting that the geometric signal will boost spatial reasoning. However, we find that simply fusing geometry features and training on standard spatial QA yields only marginal improvements on high-level multi-hop tasks. We attribute this gap to a training-signal problem: standard spatial QA can be largely answered from visual features and language priors, so the geometry pathway receives weak gradients and fails to integrate with the visual features. To provide a training signal that requires geometry, we propose \textbf{novel-view semantic rendering} as an auxiliary training task that requires the model to predict the semantic layout of an unobserved viewpoint, inspired by humans' ability to mentally simulate novel viewpoints during spatial reasoning. This task encourages joint use of both pathways: geometry provides pose-dependent visibility, while vision provides semantic content. Our auxiliary task yields consistent improvements over the geometry-augmented baseline across all three benchmarks (up to +1.6 on VSI-Bench, +2.2 on ReVSI, +2.9 on our 3D-Point-QA dataset) and our full model surpasses prior open-source methods on VSI-Bench and on ReVSI. Project page: https://yuqunw.github.io/Render2Reason/.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Wu, Y., Xiao, Y., Zou, C., Wang, S., & Hoiem, D. (2026). التصيير للاستدلال: التنبؤ الدلالي بالمنظور الجديد يحسّن الفهم المكاني في نماذج اللغة البصرية. https://omanscience.com/ar/articles/render-to-reason-novel-view-semantic-prediction-improves-spatial-understanding-in-vlms
MLA 9
Wu, Yuqun, et al. "التصيير للاستدلال: التنبؤ الدلالي بالمنظور الجديد يحسّن الفهم المكاني في نماذج اللغة البصرية." https://omanscience.com/ar/articles/render-to-reason-novel-view-semantic-prediction-improves-spatial-understanding-in-vlms.
شيكاغو (المؤلف–التاريخ)
Wu, Yuqun, Yao Xiao, Chuhang Zou, Shenlong Wang, and Derek Hoiem. 2026. "التصيير للاستدلال: التنبؤ الدلالي بالمنظور الجديد يحسّن الفهم المكاني في نماذج اللغة البصرية." https://omanscience.com/ar/articles/render-to-reason-novel-view-semantic-prediction-improves-spatial-understanding-in-vlms.
هارفارد
Wu, Y., Xiao, Y., Zou, C., Wang, S. and Hoiem, D. (2026) 'التصيير للاستدلال: التنبؤ الدلالي بالمنظور الجديد يحسّن الفهم المكاني في نماذج اللغة البصرية', Available at: https://omanscience.com/ar/articles/render-to-reason-novel-view-semantic-prediction-improves-spatial-understanding-in-vlms.
فانكوفر
Wu Y, Xiao Y, Zou C, Wang S, Hoiem D. التصيير للاستدلال: التنبؤ الدلالي بالمنظور الجديد يحسّن الفهم المكاني في نماذج اللغة البصرية. https://omanscience.com/ar/articles/render-to-reason-novel-view-semantic-prediction-improves-spatial-understanding-in-vlms
IEEE
Y. Wu, Y. Xiao, C. Zou, S. Wang, and D. Hoiem, "التصيير للاستدلال: التنبؤ الدلالي بالمنظور الجديد يحسّن الفهم المكاني في نماذج اللغة البصرية," https://omanscience.com/ar/articles/render-to-reason-novel-view-semantic-prediction-improves-spatial-understanding-in-vlms.