الباحثون

Katerina Fragkiadaki

المنشورات 5

نسخة أولية وصول مفتوح

SkillWeaver: Agentic Exploration over Neural Interaction Skills for Scalable Robot Data Generation

He Zhu, Lusen Zhao, Kwan Man Cheng وآخرون · 2026

Large-scale demonstrations have driven unprecedented progress in robot learning, yet collecting robot data through teleoperation is expensive and difficult to scale to diverse environments and long-horizon tasks. Simulation offers a scalable alternative, but existing data-generation pipelines often rely on open-loop co …

نسخة أولية وصول مفتوح

TrackEverything: Long Horizon Dense Tracking via De-Duplicating 3D Scene Representations

Existing point tracking models face a fundamental tradeoff: they can either track a sparse set of query points over long horizons, or track all points across only short clips. We introduce TrackEverything, a 3D point tracker that breaks this trade-off by representing videos as persistent 3D scene tracks in world coordi …

نسخة أولية وصول مفتوح

EmbodiedSWE: Coding Agents for Long Horizon Dexterous Robotics

Zeyu Shen, Haoxiang You, Yilang Liu وآخرون · 2026

We study coding agents for long-horizon, dexterous robotics and ask whether their solutions can provide scalable supervision for learning general robot policies. To test this, we develop EMBODIEDSWE-BENCH, a simulation benchmark for coding agents spanning contact-rich manipulation, deformable objects, and long-horizon …

نسخة أولية وصول مفتوح

HIGenNTO: Scalable Humanoid Interaction Generation via Noise-Space Trajectory Optimization

Lalit Jayanti, Kashu Yamazaki, Yuto Shibata وآخرون · 2026

Humanoid robots can acquire complex skills by imitating kinematic humanoid motion references, yet reliable references for contact-rich interactions remain difficult to obtain: motion capture deteriorates under occlusion and close physical contact, while retargeting introduces additional contact and geometric inconsiste …

نسخة أولية وصول مفتوح

GeomVLA: Unifying Scene, Motion, and Action in 3D

We present GeomVLA, a Vision-Language-Action (VLA) model that unifies perception, latent scene motion prediction, and action generation within a shared robot-centric 3D coordinate frame. Our approach lifts pretrained VLM features into spatially grounded 3D scene tokens using depth and camera calibration, while retainin …

المؤلفون المشاركون