الملخص

تمت ترجمة أجزاء من هذه الصفحة آلياً وقد تحتوي على أخطاء.

Vision-Language-Action (VLA) models acquire broad manipulation capabilities via large-scale pretraining, yet eliciting them through language requires fine-grained alignment between instructions and physical interactions. Existing robot demonstrations typically provide only coarse task descriptions, omitting how actions are executed, including which gripper acts, which object is contacted, and how it is grasped and moved. We introduce YUBI-STAG, a framework for Spatio-Temporal Annotation and Grounding that automatically enriches manipulation demonstrations with interaction-rich semantics to align pretrained VLAs with fine-grained manipulation language. Combining contact-object segmentation with vision-language models, YUBI-STAG annotates object identities, attributes and states, per-gripper actions, bimanual coordination, and spatially grounded interactions. To address YUBI-STAG's reliance on localized sequences and multi-stage VLM inference, we distill it into YUBI-VLM. YUBI-VLM directly recovers action structure and annotations from raw, unsegmented video in few inference calls and operates from wrist views alone. We evaluate both frameworks on YUBI-STAG-Bench across temporal, semantic, and spatial grounding tasks. YUBI-VLM retains much of YUBI-STAG's annotation accuracy with fewer inference calls and shorter runtime while generalizing to unseen manipulations. Finally, post-training VLA policies on these annotations aligns them with fine-grained language and contact-aware structure. Bimanual experiments demonstrate improved performance and instruction following, including control over object identity, acting gripper, target location, and spatial relations absent from original labels.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Tateno, M., Ohkawa, T., Wu, Y. H., Li, H., Matsushima, T., Sato, Y., & Ota, K. (2026). YUBI-STAG: المحاذاة الغنية بالتلامس والدلالة لنماذج VLA عبر التأريض الآلي لفيديو-لغة. https://omanscience.com/ar/articles/yubi-stag-contact-and-semantic-rich-alignment-for-vlas-via-automated-video-language-grounding

MLA 9

Tateno, Masatoshi, et al. "YUBI-STAG: المحاذاة الغنية بالتلامس والدلالة لنماذج VLA عبر التأريض الآلي لفيديو-لغة." https://omanscience.com/ar/articles/yubi-stag-contact-and-semantic-rich-alignment-for-vlas-via-automated-video-language-grounding.

شيكاغو (المؤلف–التاريخ)

Tateno, Masatoshi, Takehiko Ohkawa, Yueh-Hua Wu, Hanlong Li, Tatsuya Matsushima, Yoichi Sato, and Kei Ota. 2026. "YUBI-STAG: المحاذاة الغنية بالتلامس والدلالة لنماذج VLA عبر التأريض الآلي لفيديو-لغة." https://omanscience.com/ar/articles/yubi-stag-contact-and-semantic-rich-alignment-for-vlas-via-automated-video-language-grounding.

هارفارد

Tateno, M., Ohkawa, T., Wu, Y. H., Li, H., Matsushima, T., Sato, Y. and Ota, K. (2026) 'YUBI-STAG: المحاذاة الغنية بالتلامس والدلالة لنماذج VLA عبر التأريض الآلي لفيديو-لغة', Available at: https://omanscience.com/ar/articles/yubi-stag-contact-and-semantic-rich-alignment-for-vlas-via-automated-video-language-grounding.

فانكوفر

Tateno M, Ohkawa T, Wu YH, Li H, Matsushima T, Sato Y, et al. YUBI-STAG: المحاذاة الغنية بالتلامس والدلالة لنماذج VLA عبر التأريض الآلي لفيديو-لغة. https://omanscience.com/ar/articles/yubi-stag-contact-and-semantic-rich-alignment-for-vlas-via-automated-video-language-grounding

IEEE

M. Tateno, T. Ohkawa, Y. H. Wu, H. Li, T. Matsushima, Y. Sato, and K. Ota, "YUBI-STAG: المحاذاة الغنية بالتلامس والدلالة لنماذج VLA عبر التأريض الآلي لفيديو-لغة," https://omanscience.com/ar/articles/yubi-stag-contact-and-semantic-rich-alignment-for-vlas-via-automated-video-language-grounding.