Abstract

Vision-language-action (VLA) models bridge multimodal understanding and robotic control, advancing the dominant paradigm for embodied intelligence. However, most existing models rely on large Transformers, whose latency and energy costs hinder deployment on resource-constrained platforms. Through sparse event-driven computation, spiking neural networks offer a promising paradigm for high-performance and energy-efficient computing. Here, we propose the first Spike-driven VLA framework enabling end-to-end direct training for robotic manipulation, which mainly comprises three core components. First, we develop spiking visual and instruction encoders for multimodal perception, encoding visual observations and language instructions into sparse, reliable spike representations for subsequent cross-modal fusion. Then, we introduce Multi-Winner Spike Fusion for instruction-guided scene understanding, using bidirectional top-$k$ winner-take-all spike routing to suppress background interference and yield fused memory. Finally, we propose a Spike Action Chunking Transformer that incorporates spiking cross-attention over the fused memory and the current robot state, enabling efficient end-to-end generation of continuous action chunks for robotic control. Extensive experiments on LIBERO and Meta-World demonstrate that Spike-driven VLA achieves competitive performance with fewer parameters and lower estimated inference energy than conventional VLA models. This work establishes a foundational framework for neuromorphic VLA modeling, paving the way for future advances in resource-efficient embodied intelligence.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Wang, S., Zhang, M., Liu, M., Dai, W., Zhang, D., Zhang, J., Shan, Y., Zhou, Z., & Yang, Y. (2026). Spike-driven Vision-Language-Action Model. https://omanscience.com/en/articles/spike-driven-vision-language-action-model

MLA 9

Wang, Shuai, et al. "Spike-driven Vision-Language-Action Model." https://omanscience.com/en/articles/spike-driven-vision-language-action-model.

Chicago (author–date)

Wang, Shuai, Malu Zhang, Mingquan Liu, Weihui Dai, Dehao Zhang, Jieyuan Zhang, Yimeng Shan, Zijian Zhou, and Yang Yang. 2026. "Spike-driven Vision-Language-Action Model." https://omanscience.com/en/articles/spike-driven-vision-language-action-model.

Harvard

Wang, S., Zhang, M., Liu, M., Dai, W., Zhang, D., Zhang, J., Shan, Y., Zhou, Z. and Yang, Y. (2026) 'Spike-driven Vision-Language-Action Model', Available at: https://omanscience.com/en/articles/spike-driven-vision-language-action-model.

Vancouver

Wang S, Zhang M, Liu M, Dai W, Zhang D, Zhang J, et al. Spike-driven Vision-Language-Action Model. https://omanscience.com/en/articles/spike-driven-vision-language-action-model

IEEE

S. Wang, M. Zhang, M. Liu, W. Dai, D. Zhang, J. Zhang, Y. Shan, Z. Zhou, and Y. Yang, "Spike-driven Vision-Language-Action Model," https://omanscience.com/en/articles/spike-driven-vision-language-action-model.