نسخة أولية وصول مفتوح
VGGTWorld-VLA: Intent-Conditioned 3D World Evolution for Autonomous Driving
VGGT provides a strong foundation for geometry-centric world models by recovering unified 3D scene geometry from visual observations. Although recent extensions enable temporal 3D prediction, their future evolution remains weakly conditioned on driving intentions and actions, limiting their ability to model alternative …