Abstract

Adapting vision-language-action (VLA) models to deployment-time distribution shifts is important for reliable robotic operation, but conventional first-order adaptation can exceed the memory budget of inference-oriented deployment platforms. Zeroth-order (ZO) optimization offers a forward-only alternative with inference-level memory, but accurate gradient estimation requires many perturbation queries, making naive ZO prohibitively slow for large VLA models. We present VLA-ZO, a framework for fast ZO adaptation that exploits the structure of VLA computation. By confining adaptation to the action side, VLA-ZO keeps the expensive vision-language prefix frozen and reuses its conditioning states across perturbation queries and optimizer steps, while schedule-aware prefetching hides state-transfer overhead. On LIBERO camera-viewpoint shifts, VLA-ZO reduces end-to-end adaptation time by 25.59$\times$ at $q=16$ and 32.54$\times$ at $q=64$ relative to baseline ZO, while improving average task success from 48.27% without adaptation to 58.17% and 63.58%, respectively. These results show that making ZO faster can make larger query budgets practical, providing a promising path toward resource-efficient VLA adaptation on deployment platforms.

Keywords

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Kim, J., Kim, J., & Gong, T. (2026). VLA-ZO: Fast Zeroth-Order Adaptation for Vision-Language-Action Models. https://omanscience.com/en/articles/vla-zo-fast-zeroth-order-adaptation-for-vision-language-action-models

MLA 9

Kim, Jaemin, et al. "VLA-ZO: Fast Zeroth-Order Adaptation for Vision-Language-Action Models." https://omanscience.com/en/articles/vla-zo-fast-zeroth-order-adaptation-for-vision-language-action-models.

Chicago (author–date)

Kim, Jaemin, Jiahn Kim, and Taesik Gong. 2026. "VLA-ZO: Fast Zeroth-Order Adaptation for Vision-Language-Action Models." https://omanscience.com/en/articles/vla-zo-fast-zeroth-order-adaptation-for-vision-language-action-models.

Harvard

Kim, J., Kim, J. and Gong, T. (2026) 'VLA-ZO: Fast Zeroth-Order Adaptation for Vision-Language-Action Models', Available at: https://omanscience.com/en/articles/vla-zo-fast-zeroth-order-adaptation-for-vision-language-action-models.

Vancouver

Kim J, Kim J, Gong T. VLA-ZO: Fast Zeroth-Order Adaptation for Vision-Language-Action Models. https://omanscience.com/en/articles/vla-zo-fast-zeroth-order-adaptation-for-vision-language-action-models

IEEE

J. Kim, J. Kim, and T. Gong, "VLA-ZO: Fast Zeroth-Order Adaptation for Vision-Language-Action Models," https://omanscience.com/en/articles/vla-zo-fast-zeroth-order-adaptation-for-vision-language-action-models.