الملخص
Vision-language-action (VLA) models enable task-conditioned interaction, but extending them to scene-scale aerial manipulation remains challenging due to costly whole-body demonstrations, latency-induced action-state misalignment, and cross-site behavior composition. We present a unified framework for synthetic policy training and scene-scale execution on articulated uncrewed aerial manipulators (UAMs). A scene-reconfigurable pipeline synthesizes task-conditioned, kinodynamically feasible trajectories and synchronized multiview observations for VLA training without physical-platform demonstrations. Measured-progress-aligned realization (MPAR) aligns asynchronously returned action chunks with measured execution progress and realizes them as continuous, dynamically feasible trajectories. A relational Scene Graph grounds language goals to object instances and feasible interaction regions, while topology-guided transfer connects local behaviors across sites. Local VLA skills achieve 39/60 successes (65.0%) in simulation under oracle target and feasible-handoff conditions. Under 500-ms added latency, with and without a transient command-update stall, MPAR reduces median takeover phase error by 0.212 s over nominal-time alignment. The complete system completes 21/50 simulated multi-site missions (42.0%) and is further validated on a physical articulated UAM.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Guo, W., Jin, R., Jin, H., Xu, X., Liu, R., Zhao, H., Wang, Y., Gai, W., Cao, K., & Xie, L. (2026). From Local Whole-Body VLA Behaviors to Scene-Scale Aerial Manipulation. https://omanscience.com/ar/articles/from-local-whole-body-vla-behaviors-to-scene-scale-aerial-manipulation
MLA 9
Guo, Weixiang, et al. "From Local Whole-Body VLA Behaviors to Scene-Scale Aerial Manipulation." https://omanscience.com/ar/articles/from-local-whole-body-vla-behaviors-to-scene-scale-aerial-manipulation.
شيكاغو (المؤلف–التاريخ)
Guo, Weixiang, Rui Jin, Haotian Jin, Xinhang Xu, Ruiyang Liu, Haoran Zhao, Yi Wang, Weiqi Gai, Kun Cao, and Lihua Xie. 2026. "From Local Whole-Body VLA Behaviors to Scene-Scale Aerial Manipulation." https://omanscience.com/ar/articles/from-local-whole-body-vla-behaviors-to-scene-scale-aerial-manipulation.
هارفارد
Guo, W., Jin, R., Jin, H., Xu, X., Liu, R., Zhao, H., Wang, Y., Gai, W., Cao, K. and Xie, L. (2026) 'From Local Whole-Body VLA Behaviors to Scene-Scale Aerial Manipulation', Available at: https://omanscience.com/ar/articles/from-local-whole-body-vla-behaviors-to-scene-scale-aerial-manipulation.
فانكوفر
Guo W, Jin R, Jin H, Xu X, Liu R, Zhao H, et al. From Local Whole-Body VLA Behaviors to Scene-Scale Aerial Manipulation. https://omanscience.com/ar/articles/from-local-whole-body-vla-behaviors-to-scene-scale-aerial-manipulation
IEEE
W. Guo, R. Jin, H. Jin, X. Xu, R. Liu, H. Zhao, Y. Wang, W. Gai, K. Cao, and L. Xie, "From Local Whole-Body VLA Behaviors to Scene-Scale Aerial Manipulation," https://omanscience.com/ar/articles/from-local-whole-body-vla-behaviors-to-scene-scale-aerial-manipulation.