Abstract
Three-dimensional perception is critical for robotic manipulation, particularly for high-precision tasks, as recovering metric depth and precise 3D object positions from monocular RGB observations is inherently ill-posed. However, many Vision-Language-Action (VLA) models rely solely on RGB observations for perception. Leveraging recent advances in foundation models for stereo matching, we introduce ExStereo, a stereo module that augments pre-trained 2D VLAs with 3D perception. ExStereo reconstructs scene geometry from stereo image pairs and renders multi-view observations as an explicit stereo representation for stereo feature extraction. The action tokens from the action expert selectively attend to the resulting stereo tokens through our proposed action-stereo cross-attention mechanism, enabling the policy to generate robot actions conditioned on 3D scene information. To learn robust 3D representations, we introduce a mid-training stage before task-specific post-training, using a self-supervised learning objective on large-scale stereo data. We validate our approach by fine-tuning two publicly available VLAs, $π_{0.5}$ and SmolVLA, and evaluate them in simulation and on a real-world bimanual PiPER platform. Across both settings, VLAs fine-tuned with ExStereo consistently outperform baselines, demonstrating the effectiveness of stereo perception for robotic manipulation. Our project website is at: https://exstereo-vla.github.io/ExStereo/.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Liu, I. C. A., Chen, J., Sukhatme, G. S., & Seita, D. (2026). ExStereo: Lifting 2D Vision-Language-Action Models to 3D with Explicit Stereo Representations. https://omanscience.com/en/articles/exstereo-lifting-2d-vision-language-action-models-to-3d-with-explicit-stereo-representations
MLA 9
Liu, I-Chun Arthur, et al. "ExStereo: Lifting 2D Vision-Language-Action Models to 3D with Explicit Stereo Representations." https://omanscience.com/en/articles/exstereo-lifting-2d-vision-language-action-models-to-3d-with-explicit-stereo-representations.
Chicago (author–date)
Liu, I-Chun Arthur, Jason Chen, Gaurav S. Sukhatme, and Daniel Seita. 2026. "ExStereo: Lifting 2D Vision-Language-Action Models to 3D with Explicit Stereo Representations." https://omanscience.com/en/articles/exstereo-lifting-2d-vision-language-action-models-to-3d-with-explicit-stereo-representations.
Harvard
Liu, I. C. A., Chen, J., Sukhatme, G. S. and Seita, D. (2026) 'ExStereo: Lifting 2D Vision-Language-Action Models to 3D with Explicit Stereo Representations', Available at: https://omanscience.com/en/articles/exstereo-lifting-2d-vision-language-action-models-to-3d-with-explicit-stereo-representations.
Vancouver
Liu ICA, Chen J, Sukhatme GS, Seita D. ExStereo: Lifting 2D Vision-Language-Action Models to 3D with Explicit Stereo Representations. https://omanscience.com/en/articles/exstereo-lifting-2d-vision-language-action-models-to-3d-with-explicit-stereo-representations
IEEE
I. C. A. Liu, J. Chen, G. S. Sukhatme, and D. Seita, "ExStereo: Lifting 2D Vision-Language-Action Models to 3D with Explicit Stereo Representations," https://omanscience.com/en/articles/exstereo-lifting-2d-vision-language-action-models-to-3d-with-explicit-stereo-representations.