Abstract

Vision-Language-Action (VLA) models enable generalizable robotic control but remain computationally expensive. Token caching provides a training-free, plug-and-play acceleration alternative. However, existing VLA caching does not fully exploit a key inductive bias of VLA models: text-vision synergy, wherein textual semantics guide the precise visual grounding of task-relevant regions. In particular, existing designs insufficiently account for head-wise reliability in attention aggregation and layer-wise stability in cache reuse. To address this, we propose Text-Vision Synergistic Token Caching (TVCache), a training-free framework for efficient VLA inference. TVCache filters attention heads based on text-vision information focus to improve task-relevant and physically consistent visual grounding. Concurrently, we introduce a reuse-layer selection mechanism guided by text-vision entropy differences to avoid caching unstable representations and improve cache resource allocation. Extensive experiments across four representative VLA models, two simulation benchmarks, and real-world robotic tasks demonstrate the effectiveness and generality of TVCache. At matched token-retention ratios, TVCache consistently improves task success over existing VLA caching with comparable computational cost. On OpenVLA-OFT, it improves average success by up to 14.5 percentage points over VLA-Cache at 12.5% retention while reducing FLOPs by 2.45x relative to full-token inference.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Li, Q., Zhang, C., Chen, J., Tong, Z., Zhang, J., & Zhang, H. (2026). Text-Vision Synergistic Token Caching: A Training-Free Framework for Efficient Vision-Language-Action Inference. https://omanscience.com/en/articles/text-vision-synergistic-token-caching-a-training-free-framework-for-efficient-vision-language-action-inference

MLA 9

Li, Qianer, et al. "Text-Vision Synergistic Token Caching: A Training-Free Framework for Efficient Vision-Language-Action Inference." https://omanscience.com/en/articles/text-vision-synergistic-token-caching-a-training-free-framework-for-efficient-vision-language-action-inference.

Chicago (author–date)

Li, Qianer, Chengjie Zhang, Jingwen Chen, Zanjia Tong, Jiyuan Zhang, and Hong Zhang. 2026. "Text-Vision Synergistic Token Caching: A Training-Free Framework for Efficient Vision-Language-Action Inference." https://omanscience.com/en/articles/text-vision-synergistic-token-caching-a-training-free-framework-for-efficient-vision-language-action-inference.

Harvard

Li, Q., Zhang, C., Chen, J., Tong, Z., Zhang, J. and Zhang, H. (2026) 'Text-Vision Synergistic Token Caching: A Training-Free Framework for Efficient Vision-Language-Action Inference', Available at: https://omanscience.com/en/articles/text-vision-synergistic-token-caching-a-training-free-framework-for-efficient-vision-language-action-inference.

Vancouver

Li Q, Zhang C, Chen J, Tong Z, Zhang J, Zhang H. Text-Vision Synergistic Token Caching: A Training-Free Framework for Efficient Vision-Language-Action Inference. https://omanscience.com/en/articles/text-vision-synergistic-token-caching-a-training-free-framework-for-efficient-vision-language-action-inference

IEEE

Q. Li, C. Zhang, J. Chen, Z. Tong, J. Zhang, and H. Zhang, "Text-Vision Synergistic Token Caching: A Training-Free Framework for Efficient Vision-Language-Action Inference," https://omanscience.com/en/articles/text-vision-synergistic-token-caching-a-training-free-framework-for-efficient-vision-language-action-inference.