نسخة أولية وصول مفتوح
ExStereo: Lifting 2D Vision-Language-Action Models to 3D with Explicit Stereo Representations
Three-dimensional perception is critical for robotic manipulation, particularly for high-precision tasks, as recovering metric depth and precise 3D object positions from monocular RGB observations is inherently ill-posed. However, many Vision-Language-Action (VLA) models rely solely on RGB observations for perception. …