Preprint Open access
When and What to Prune? Stage-Aware Visual Token Pruning for Efficient VLA
Visual token pruning is an effective way to accelerate vision-language models and is especially useful for vision-language-action (VLA) inference, where many visual tokens must be processed before predicting robot actions. Existing pruning methods usually estimate which tokens can be pruned based on attention scores or …