الملخص

تمت ترجمة أجزاء من هذه الصفحة آلياً وقد تحتوي على أخطاء.

Vision-language-action (VLA) policies can struggle with manipulation tasks with complex obstacle geometries due to partial observability. These complex geometries can lead to similar visual observations or robot configurations requiring qualitatively different actions, a distinction that can be quantified using topological signatures. While motion planners with full knowledge of environment geometries and object states can reason about these signatures in planning, this information is often not known at deployment. To address this issue, we present a topology-guided visual-prompting framework that uses simulation-based planning to augment a nominal demonstration dataset and provides vision-based guidance at deployment. Our method uses a Gauss-Linking-Integral topological signature representation to capture important topological properties of the environment. Using privileged geometry information from a simulation approximation of our environment, we augment a VLA fine-tuning dataset with trajectories that move the system to a demonstrated signature and, from the new configuration, resume task execution. A vision-language model (VLM) is fine-tuned on the same dataset to both predict signatures from live camera observations and predict end-effector waypoints, which are rendered as visual prompts on the observations to guide the VLA. Across three simulated bimanual tasks and a real-world box pickup task, our method outperforms a VLA fine-tuned only on nominal demonstrations and a VLM-prompting baseline that can remove topology-relevant information from observations. On hardware, it exceeds the strongest baseline by 40% in task success. Project website: https://topology-vla.github.io.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Wu, H., Kumar, A., & Berenson, D. (2026). تحفيز بصري مدرك للطوبولوجيا لسياسات الرؤية واللغة والفعل. https://omanscience.com/ar/articles/topology-informed-visual-prompting-for-vision-language-action-policies

MLA 9

Wu, Haoyang, et al. "تحفيز بصري مدرك للطوبولوجيا لسياسات الرؤية واللغة والفعل." https://omanscience.com/ar/articles/topology-informed-visual-prompting-for-vision-language-action-policies.

شيكاغو (المؤلف–التاريخ)

Wu, Haoyang, Abhinav Kumar, and Dmitry Berenson. 2026. "تحفيز بصري مدرك للطوبولوجيا لسياسات الرؤية واللغة والفعل." https://omanscience.com/ar/articles/topology-informed-visual-prompting-for-vision-language-action-policies.

هارفارد

Wu, H., Kumar, A. and Berenson, D. (2026) 'تحفيز بصري مدرك للطوبولوجيا لسياسات الرؤية واللغة والفعل', Available at: https://omanscience.com/ar/articles/topology-informed-visual-prompting-for-vision-language-action-policies.

فانكوفر

Wu H, Kumar A, Berenson D. تحفيز بصري مدرك للطوبولوجيا لسياسات الرؤية واللغة والفعل. https://omanscience.com/ar/articles/topology-informed-visual-prompting-for-vision-language-action-policies

IEEE

H. Wu, A. Kumar, and D. Berenson, "تحفيز بصري مدرك للطوبولوجيا لسياسات الرؤية واللغة والفعل," https://omanscience.com/ar/articles/topology-informed-visual-prompting-for-vision-language-action-policies.