Abstract

Vision-language-action (VLA) models provide broad, instruction-conditioned manipulation behaviors, but their physical execution can remain imprecise during contact-rich interaction. Residual reinforcement learning (RL) can correct such errors while keeping the VLA frozen, but real-robot RL is costly and safety-critical. We propose VLA Latent-Conditioned RL (VLaRL), which enables residual RL for frozen VLAs to be trained in simulation and deployed on real robots without real-world RL or online adaptation. The key challenge is transferring the learned residual policy despite the visual gap between simulation and reality. Rather than requiring pixel-level visual correspondence, VLaRL uses the VLA's internal vision-language latent representation to condition residual control and as the sim-to-real transfer interface, and learns a lightweight mapper that transforms simulation-derived latents toward the real latent distribution. Across four contact-rich manipulation tasks and two VLA backbones, VLaRL improves real-world success in all task-backbone combinations, while controlled ablations demonstrate the importance of both latent conditioning and latent alignment for transferring simulation-trained residual control.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Saito, N., Kim, K., Kim, H., Ikeuchi, K., & Matsushita, Y. (2026). VLaRL: Augmenting Vision-Language-Action Models with Simulation-Trained Latent-Conditioned Residual RL. https://omanscience.com/en/articles/vlarl-augmenting-vision-language-action-models-with-simulation-trained-latent-conditioned-residual-rl

MLA 9

Saito, Namiko, et al. "VLaRL: Augmenting Vision-Language-Action Models with Simulation-Trained Latent-Conditioned Residual RL." https://omanscience.com/en/articles/vlarl-augmenting-vision-language-action-models-with-simulation-trained-latent-conditioned-residual-rl.

Chicago (author–date)

Saito, Namiko, Kinam Kim, Heecheol Kim, Katsushi Ikeuchi, and Yasuyuki Matsushita. 2026. "VLaRL: Augmenting Vision-Language-Action Models with Simulation-Trained Latent-Conditioned Residual RL." https://omanscience.com/en/articles/vlarl-augmenting-vision-language-action-models-with-simulation-trained-latent-conditioned-residual-rl.

Harvard

Saito, N., Kim, K., Kim, H., Ikeuchi, K. and Matsushita, Y. (2026) 'VLaRL: Augmenting Vision-Language-Action Models with Simulation-Trained Latent-Conditioned Residual RL', Available at: https://omanscience.com/en/articles/vlarl-augmenting-vision-language-action-models-with-simulation-trained-latent-conditioned-residual-rl.

Vancouver

Saito N, Kim K, Kim H, Ikeuchi K, Matsushita Y. VLaRL: Augmenting Vision-Language-Action Models with Simulation-Trained Latent-Conditioned Residual RL. https://omanscience.com/en/articles/vlarl-augmenting-vision-language-action-models-with-simulation-trained-latent-conditioned-residual-rl

IEEE

N. Saito, K. Kim, H. Kim, K. Ikeuchi, and Y. Matsushita, "VLaRL: Augmenting Vision-Language-Action Models with Simulation-Trained Latent-Conditioned Residual RL," https://omanscience.com/en/articles/vlarl-augmenting-vision-language-action-models-with-simulation-trained-latent-conditioned-residual-rl.