الملخص
Grounding is a core capability of spatial vision-language models, yet most existing work focuses only on where a referred object is. Many 3D tasks also require knowing how it is oriented. Although existing 3D VLMs may predict oriented boxes, box pose does not explicitly capture object-centric orientation or symmetry-induced ambiguities. We introduce orientation grounding, a referring grounding task that predicts an object's 6D orientation and axial symmetry from a language or box query in single-view or multi-view scenes. To support this task, we construct ReferOri, with 331K multi-view and 387K single-view orientation-grounding queries obtained through scalable reconstruction, consistency checking, and human verification. We further present OG-VLM, which adapts a 3D VLM with structured box/orientation outputs, sign and symmetry tokens, and geometry-aware auxiliary losses. Across single-view and multi-view benchmarks, OG-VLM substantially outperforms orientation-aware VLM baselines and surpasses object-level orientation foundation models on scene-level referring benchmarks, showing that explicit orientation grounding is a distinct and learnable capability beyond localization. Downstream results validate its benefit for orientation-related spatial reasoning.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Liang, T., Liu, D., Wang, N., Chaudhary, V., & Yin, Y. (2026). Toward Comprehensive 3D Grounding: Orientation Grounding through Vision-Language Models. https://omanscience.com/ar/articles/toward-comprehensive-3d-grounding-orientation-grounding-through-vision-language-models
MLA 9
Liang, Tuo, et al. "Toward Comprehensive 3D Grounding: Orientation Grounding through Vision-Language Models." https://omanscience.com/ar/articles/toward-comprehensive-3d-grounding-orientation-grounding-through-vision-language-models.
شيكاغو (المؤلف–التاريخ)
Liang, Tuo, Disheng Liu, Nengbo Wang, Vipin Chaudhary, and Yu Yin. 2026. "Toward Comprehensive 3D Grounding: Orientation Grounding through Vision-Language Models." https://omanscience.com/ar/articles/toward-comprehensive-3d-grounding-orientation-grounding-through-vision-language-models.
هارفارد
Liang, T., Liu, D., Wang, N., Chaudhary, V. and Yin, Y. (2026) 'Toward Comprehensive 3D Grounding: Orientation Grounding through Vision-Language Models', Available at: https://omanscience.com/ar/articles/toward-comprehensive-3d-grounding-orientation-grounding-through-vision-language-models.
فانكوفر
Liang T, Liu D, Wang N, Chaudhary V, Yin Y. Toward Comprehensive 3D Grounding: Orientation Grounding through Vision-Language Models. https://omanscience.com/ar/articles/toward-comprehensive-3d-grounding-orientation-grounding-through-vision-language-models
IEEE
T. Liang, D. Liu, N. Wang, V. Chaudhary, and Y. Yin, "Toward Comprehensive 3D Grounding: Orientation Grounding through Vision-Language Models," https://omanscience.com/ar/articles/toward-comprehensive-3d-grounding-orientation-grounding-through-vision-language-models.