الباحثون

An Liu

المنشورات 1

نسخة أولية وصول مفتوح

GeoScaffold: Learning Compact Geometric Latents via Reconstruction for Efficient Vision-Language Navigation

Yixuan Jiang, Wentong Li, An Liu وآخرون · 2026

Recent vision-and-language navigation (VLN) systems increasingly adopt streaming Video-LLM policies that map egocentric RGB observations and instructions directly to low-level actions. Yet these policies inherit weak 3D geometric priors from 2D pretraining. Existing geometry-aware extensions charge a persistent inferen …

المؤلفون المشاركون