الباحثون

Jitendra Malik

المنشورات 5

نسخة أولية وصول مفتوح

EyeRobot 2.0: Active Gaze for Precise Manipulation without Wrist Cameras

Kush Hari, Justin Kerr, Nidhya Shivakumar وآخرون · 2026

Inspired by human vision, we introduce a framework using active gaze to enable fine-grained bimanual manipulation with only a single stereo camera. EyeRobot 2.0 physically attends to a 3D fixation point in the scene by swiveling two eye viewpoints to center their gaze on it. The resulting images are processed foveally …

نسخة أولية وصول مفتوح

OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?

Jiaxin Ge, Yiming Qin, Ji Xie وآخرون · 2026

Training a model to generate visual content can encourage it to learn rich perceptual capabilities related to geometry, spatial relationships, and objectness; yet, its benefits for visual understanding remain unclear. We ask: when and how does visual generation supervision improve visual understanding? We study control …

نسخة أولية وصول مفتوح

Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy

Tara Sadjadpour, Siming He, C. K. Wolfe وآخرون · 2026

Human hand-object interactions (HOIs) provide a rich source of demonstrations for dexterous manipulation, but learning directly from them presents challenges in bridging morphology gaps, ensuring dynamical feasibility, and sim-to-real deployment. We present Morphometric Imitation, a three-stage framework that transform …

نسخة أولية وصول مفتوح

A Scene Language Model for Open-Vocabulary Scene Mapping

Adam Lilja, Fabio Hübel, Siming He وآخرون · 2026

Open-vocabulary 3D scene mapping aims to build a persistent representation of the objects in an environment. Existing systems typically rely on engineered mapping pipelines to associate observations, merge information across views, and maintain a consistent scene representation over time. Many additionally store featur …

المؤلفون المشاركون