الباحثون

Kangmin Kim

المنشورات 2

نسخة أولية وصول مفتوح

GeoBridge-VLA: Geometry-Aware Residual Adaptation for Vision-Language-Action Models

Hyun Song, Kangmin Kim, Loren Jinsoo Um وآخرون · 2026

Vision-language-action (VLA) models encode semantic information from vision-language pretraining, but manipulation also requires precise spatial reasoning. We present GeoBridge-VLA, a two-stage method for learning geometric features from a pretrained VLA's frozen visual encoder and using them for action prediction. Sta …

نسخة أولية وصول مفتوح

SPACE-CLIPv2: Decoding Local Geometry from Frozen CLIP for Monocular Depth Estimation

Hyun Song, Taewan Cho, Kangmin Kim وآخرون · 2026

Vision-language foundation models such as CLIP provide strong semantic representations, but their patch tokens are not directly optimized for dense metric geometry. SPACE-CLIP showed that frozen CLIP features can support monocular depth estimation through layer-group feature fusion, yet it leaves open how neighboring C …

المؤلفون المشاركون