الباحثون

Shriram Damodaran

المنشورات 2

نسخة أولية وصول مفتوح

ActionGround: Training-Free Runtime Refinement of Frozen VLA Policies

Vision-Language-Action (VLA) models map visual observations and language instructions directly to robot actions, but they do not explicitly represent the phase structure of manipulation tasks or the rigid-body dynamics governing execution. We present ActionGround, a neuro-symbolic, training-free runtime layer that wrap …

نسخة أولية وصول مفتوح

SphMind: Towards Robust, Training-Free VLM-based Spatial Reasoning with a 360 Camera

Omnidirectional or 360 cameras provide embodied AI agents with a holistic, wide field-of-view (FoV) view of their surroundings, motivating the use of Multi-modal Large Language Models (MLLMs) for omnidirectional spatial reasoning. However, most MLLMs are trained on conventional 2D perspective images and struggle with t …

المؤلفون المشاركون