الملخص

Recent advances in 3D reconstruction have progressed from per-scene optimization to feed-forward inference, and semantic scene understanding has followed suit -- yet existing methods remain confined to object-centric perception, neglecting spatial relations between objects. We formulate 3D spatial relation segmentation in a feed-forward, pose-free multi-view setting: given a visually specified subject and a relational text query, the model segments the target across views without receiving its category name. To this end, we propose RelationVGGT, a novel feed-forward framework that integrates semantic features from a visual foundation model with geometry-aware representations from a 3D geometry foundation model and leverages a relation transformer for subject-conditioned, cross-view relation prediction -- requiring neither per-scene optimization nor known camera poses. We additionally provide a fully automated annotation pipeline built on ScanNet++ with VLMs and LLMs, enabling scalable training data generation for this new task.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Kim, M., Choe, J., Lee, J., Wang, Y. C. F., & Kim, S. J. (2026). RelationVGGT: Visual Geometry Transformers for 3D Spatial Relation Segmentation. https://omanscience.com/ar/articles/relationvggt-visual-geometry-transformers-for-3d-spatial-relation-segmentation

MLA 9

Kim, Minsu, et al. "RelationVGGT: Visual Geometry Transformers for 3D Spatial Relation Segmentation." https://omanscience.com/ar/articles/relationvggt-visual-geometry-transformers-for-3d-spatial-relation-segmentation.

شيكاغو (المؤلف–التاريخ)

Kim, Minsu, Jaesung Choe, Jiwoo Lee, Yu-Chiang Frank Wang, and Seon Joo Kim. 2026. "RelationVGGT: Visual Geometry Transformers for 3D Spatial Relation Segmentation." https://omanscience.com/ar/articles/relationvggt-visual-geometry-transformers-for-3d-spatial-relation-segmentation.

هارفارد

Kim, M., Choe, J., Lee, J., Wang, Y. C. F. and Kim, S. J. (2026) 'RelationVGGT: Visual Geometry Transformers for 3D Spatial Relation Segmentation', Available at: https://omanscience.com/ar/articles/relationvggt-visual-geometry-transformers-for-3d-spatial-relation-segmentation.

فانكوفر

Kim M, Choe J, Lee J, Wang YCF, Kim SJ. RelationVGGT: Visual Geometry Transformers for 3D Spatial Relation Segmentation. https://omanscience.com/ar/articles/relationvggt-visual-geometry-transformers-for-3d-spatial-relation-segmentation

IEEE

M. Kim, J. Choe, J. Lee, Y. C. F. Wang, and S. J. Kim, "RelationVGGT: Visual Geometry Transformers for 3D Spatial Relation Segmentation," https://omanscience.com/ar/articles/relationvggt-visual-geometry-transformers-for-3d-spatial-relation-segmentation.