Preprint Open access
GeoBridge-VLA: Geometry-Aware Residual Adaptation for Vision-Language-Action Models
Vision-language-action (VLA) models encode semantic information from vision-language pretraining, but manipulation also requires precise spatial reasoning. We present GeoBridge-VLA, a two-stage method for learning geometric features from a pretrained VLA's frozen visual encoder and using them for action prediction. Sta …