الملخص

Vision-language-action (VLA) models built on pretrained vision-language models have demonstrated strong performance across diverse robotic manipulation tasks. However, VLA models that directly map current 2D observations to actions often lack sufficient spatial and temporal understanding, limiting their performance in precise and long-horizon manipulation. Recent methods enhance VLA models through geometric supervision and future-state prediction across the entire scene. However, these methods can suffer from redundant scene information, distracting the model from learning the geometry and dynamics relevant to the current interaction. To address this issue, we propose FOCAL-VLA, a framework that combines subtask-guided geometry distillation with implicit world modeling to learn representations of current spatial structure and future interaction dynamics. To focus geometric learning on the current subtask, we transfer geometric knowledge from VGGT to the VLA model by aligning geometry latents with features from subtask-relevant image regions. To capture the future 3D evolution of the current interaction, we incorporate implicit world modeling using Track4World features from current and future demonstration frames. The two complementary representations jointly guide action generation without running VGGT or Track4World at inference time. Experiments show that FOCAL-VLA outperforms baselines on both simulation benchmarks and real-world manipulation tasks. Project website: https://zhiyuan-gao.github.io/FOCAL-VLA/.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Gao, Z., Wen, D., Zhan, Y., Khoshnazar, M., Schäfer, J., Peng, K., & Beetz, M. (2026). FOCAL-VLA: Subtask-Guided Geometry Distillation and Implicit World Modeling for Vision-Language-Action Models. https://omanscience.com/ar/articles/focal-vla-subtask-guided-geometry-distillation-and-implicit-world-modeling-for-vision-language-action-models

MLA 9

Gao, Zhiyuan, et al. "FOCAL-VLA: Subtask-Guided Geometry Distillation and Implicit World Modeling for Vision-Language-Action Models." https://omanscience.com/ar/articles/focal-vla-subtask-guided-geometry-distillation-and-implicit-world-modeling-for-vision-language-action-models.

شيكاغو (المؤلف–التاريخ)

Gao, Zhiyuan, Di Wen, Yanxiang Zhan, Mohammad Khoshnazar, Jeroen Schäfer, Kunyu Peng, and Michael Beetz. 2026. "FOCAL-VLA: Subtask-Guided Geometry Distillation and Implicit World Modeling for Vision-Language-Action Models." https://omanscience.com/ar/articles/focal-vla-subtask-guided-geometry-distillation-and-implicit-world-modeling-for-vision-language-action-models.

هارفارد

Gao, Z., Wen, D., Zhan, Y., Khoshnazar, M., Schäfer, J., Peng, K. and Beetz, M. (2026) 'FOCAL-VLA: Subtask-Guided Geometry Distillation and Implicit World Modeling for Vision-Language-Action Models', Available at: https://omanscience.com/ar/articles/focal-vla-subtask-guided-geometry-distillation-and-implicit-world-modeling-for-vision-language-action-models.

فانكوفر

Gao Z, Wen D, Zhan Y, Khoshnazar M, Schäfer J, Peng K, et al. FOCAL-VLA: Subtask-Guided Geometry Distillation and Implicit World Modeling for Vision-Language-Action Models. https://omanscience.com/ar/articles/focal-vla-subtask-guided-geometry-distillation-and-implicit-world-modeling-for-vision-language-action-models

IEEE

Z. Gao, D. Wen, Y. Zhan, M. Khoshnazar, J. Schäfer, K. Peng, and M. Beetz, "FOCAL-VLA: Subtask-Guided Geometry Distillation and Implicit World Modeling for Vision-Language-Action Models," https://omanscience.com/ar/articles/focal-vla-subtask-guided-geometry-distillation-and-implicit-world-modeling-for-vision-language-action-models.