الملخص
Vision-Language-Action (VLA) models achieve strong robotic manipulation performance but incur high computational costs from processing long token sequences at every control step, limiting real-time deployment. Visual token pruning offers a direct solution, as visual patches dominate the input sequence and contain considerable redundancy. Existing approaches, however, either rely on indirect training-free heuristics, such as attention scores and motion thresholds, or require costly fine-tuning of the base VLA model. We introduce VLA-ACL (Action Consistency Learning), which learns a lightweight visual token pruning policy through action-level supervision while keeping the base VLA model entirely frozen. The training objective encourages actions produced from pruned visual contexts to remain consistent with the full-context teacher, with ground-truth actions as auxiliary supervision. This directly ties token selection to its effect on the downstream control output. Experiments on LIBERO and real-world manipulation tasks show that VLA-ACL prunes up to 87.5% of visual tokens while retaining competitive performance, reduces computation by up to 75%, and achieves a 1.5x inference speedup. These results establish a stronger performance-efficiency trade-off than existing frozen-VLA pruning methods and demonstrate the value of action-level supervision for visual token selection. Code is available at https://github.com/du-owen/VLA-ACL.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Du, O., Yue, Y., Zhang, J., Pi, J., Chen, C. B., & Huang, G. (2026). VLA-ACL: Action-Consistent Visual Token Pruning for Efficient Vision-Language-Action Models. https://omanscience.com/ar/articles/vla-acl-action-consistent-visual-token-pruning-for-efficient-vision-language-action-models
MLA 9
Du, Owen, et al. "VLA-ACL: Action-Consistent Visual Token Pruning for Efficient Vision-Language-Action Models." https://omanscience.com/ar/articles/vla-acl-action-consistent-visual-token-pruning-for-efficient-vision-language-action-models.
شيكاغو (المؤلف–التاريخ)
Du, Owen, Yang Yue, Jie Zhang, Jiaqi Pi, Chi Bene Chen, and Gao Huang. 2026. "VLA-ACL: Action-Consistent Visual Token Pruning for Efficient Vision-Language-Action Models." https://omanscience.com/ar/articles/vla-acl-action-consistent-visual-token-pruning-for-efficient-vision-language-action-models.
هارفارد
Du, O., Yue, Y., Zhang, J., Pi, J., Chen, C. B. and Huang, G. (2026) 'VLA-ACL: Action-Consistent Visual Token Pruning for Efficient Vision-Language-Action Models', Available at: https://omanscience.com/ar/articles/vla-acl-action-consistent-visual-token-pruning-for-efficient-vision-language-action-models.
فانكوفر
Du O, Yue Y, Zhang J, Pi J, Chen CB, Huang G. VLA-ACL: Action-Consistent Visual Token Pruning for Efficient Vision-Language-Action Models. https://omanscience.com/ar/articles/vla-acl-action-consistent-visual-token-pruning-for-efficient-vision-language-action-models
IEEE
O. Du, Y. Yue, J. Zhang, J. Pi, C. B. Chen, and G. Huang, "VLA-ACL: Action-Consistent Visual Token Pruning for Efficient Vision-Language-Action Models," https://omanscience.com/ar/articles/vla-acl-action-consistent-visual-token-pruning-for-efficient-vision-language-action-models.