الملخص

تمت ترجمة أجزاء من هذه الصفحة آلياً وقد تحتوي على أخطاء.

Visual token pruning is an effective way to accelerate vision-language models and is especially useful for vision-language-action (VLA) inference, where many visual tokens must be processed before predicting robot actions. Existing pruning methods usually estimate which tokens can be pruned based on attention scores or feature diversity, retaining tokens that are either highly attended or visually different from others. However, most of them use fixed pruning schedules, such as pruning once at a preset layer or pruning at uniformly spaced layers. Such schedules can be risky for VLA models, because the model may not know which visual regions matter for the action in early layers. Tokens that look unimportant at first may become useful after the model combines visual observations with the language instruction. In this work, we propose SAPrune, a training-free visual token pruning framework for efficient VLA inference. Instead of pruning at fixed layers, SAPrune uses a small calibration set to observe how action-to-visual attention changes across layers, and chooses pruning layers only after the attention pattern becomes more reliable. At each selected layer, SAPrune applies a dual-path pruning rule: one path protects strongly attended visual tokens from pruning, while the other prevents useful surrounding context from being discarded. Experiments on LIBERO, SIMPLER, and real-world robotic tasks show that SAPrune prunes 87.5% of visual tokens and achieves up to 1.718x inference speedup while maintaining competitive task success rates.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Shi, T., Xiong, H., Gong, Z., Lu, Q., & Li, L. (2026). متى وماذا نقلّم؟ تقليم الرموز البصرية المدرك للمرحلة لنماذج VLA الفعّالة. https://omanscience.com/ar/articles/when-and-what-to-prune-stage-aware-visual-token-pruning-for-efficient-vla

MLA 9

Shi, Tianjun, et al. "متى وماذا نقلّم؟ تقليم الرموز البصرية المدرك للمرحلة لنماذج VLA الفعّالة." https://omanscience.com/ar/articles/when-and-what-to-prune-stage-aware-visual-token-pruning-for-efficient-vla.

شيكاغو (المؤلف–التاريخ)

Shi, Tianjun, Haotian Xiong, Ziyu Gong, Qi Lu, and Li Li. 2026. "متى وماذا نقلّم؟ تقليم الرموز البصرية المدرك للمرحلة لنماذج VLA الفعّالة." https://omanscience.com/ar/articles/when-and-what-to-prune-stage-aware-visual-token-pruning-for-efficient-vla.

هارفارد

Shi, T., Xiong, H., Gong, Z., Lu, Q. and Li, L. (2026) 'متى وماذا نقلّم؟ تقليم الرموز البصرية المدرك للمرحلة لنماذج VLA الفعّالة', Available at: https://omanscience.com/ar/articles/when-and-what-to-prune-stage-aware-visual-token-pruning-for-efficient-vla.

فانكوفر

Shi T, Xiong H, Gong Z, Lu Q, Li L. متى وماذا نقلّم؟ تقليم الرموز البصرية المدرك للمرحلة لنماذج VLA الفعّالة. https://omanscience.com/ar/articles/when-and-what-to-prune-stage-aware-visual-token-pruning-for-efficient-vla

IEEE

T. Shi, H. Xiong, Z. Gong, Q. Lu, and L. Li, "متى وماذا نقلّم؟ تقليم الرموز البصرية المدرك للمرحلة لنماذج VLA الفعّالة," https://omanscience.com/ar/articles/when-and-what-to-prune-stage-aware-visual-token-pruning-for-efficient-vla.