الباحثون

Jiajun Wu

المنشورات 13

نسخة أولية وصول مفتوح

BrickBench: Evaluating Agentic Brick Design

Peter Kulits, Yiqing Xu, R. Kenny Jones وآخرون · 2026

We propose BrickBench, a benchmark for agentic text-conditioned LEGO-set design. Given a prompt, an agent is tasked with producing an assembly that not only satisfies semantic and design criteria, but that can also be physically built. To do so, it must select parts from a discrete library and reason jointly about loca …

نسخة أولية وصول مفتوح

Cross-Embodiment Robot Foundation World Models with Latent Actions

The diversity of robot embodiments and action spaces makes it challenging to build robot world models that generalize across different embodiments. We introduce the Latent Action-Conditioned Robot World Model (LAC-WM), which operates within a learned unified latent action space shared across diverse embodiments. This u …

نسخة أولية وصول مفتوح

Co-Evolving Robot Orchestrators and Policies through Deployment

Xilun Zhang, Maggie Wang, Erik Bauer وآخرون · 2026

Vision-language-action (VLA) policies trained on large datasets are capable within their training domains, yet they still fail to generalize to the variety of situations a robot meets in real-world deployment. Agentic robot systems complement the policy with a vision-language model (VLM) orchestrator that learns when t …

نسخة أولية وصول مفتوح

OpenWAM: An Open Framework for Composable World-Action Models

Heng Yu, David D. Yuan, Juze Zhang وآخرون · 2026

World-action models (WAMs) couple future prediction with robot control, yet existing systems often vary the video backbone, interaction structure, supervision, and inference procedure simultaneously, making their design choices difficult to compare. We introduce OPENWAM, an open world-action modeling framework built ar …

نسخة أولية وصول مفتوح

MobileVISTA: Generative Data Augmentation for Pose Generalization in Mobile Manipulation

Mobile manipulators such as humanoid robots are increasingly deployed in dynamic, unstructured environments to perform dexterous manipulation tasks. However, end-to-end manipulation policies trained to imitate demonstration data collected from a single robot pose are brittle: even centimeter-scale deviations in robot p …

نسخة أولية وصول مفتوح

Robot Learning with Visual Predicted Force

Haonan Chen, Feiyang Wu, Yuxiang Ma وآخرون · 2026

Force-aware manipulation typically relies on specialized force or tactile sensors. We show that force-aware manipulation can instead be achieved through visual force prediction from the deformation of a compliant Fin Ray gripper. Our approach trains two models. First, we train a visual force estimator on calibration da …

نسخة أولية وصول مفتوح

4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes

Ruihong Shen, Žiga Kovačič, Peter Kulits وآخرون · 2026

We introduce 4DCodeBench, a benchmark for 4D inverse graphics through code generation, in which agents reconstruct dynamic scenes from video as executable graphics programs. To accomplish this, agents must translate visual observations into compact representations of scene structure and dynamics, by implementing abstra …

نسخة أولية وصول مفتوح

DITTO-X: Forward and Reverse Teleoperation for Dexterous Manipulation and Human Intervention

Zhanpeng He, Joaquin Palacios, Zhangyu Wang وآخرون · 2026

Teleoperated demonstrations are a primary source of data for robot manipulation, and teleoperated interventions are a primary mechanism for correcting policies at deployment. Yet most teleoperation systems close the loop through vision alone and are built around parallel-jaw grippers, limiting both what the robot can e …

نسخة أولية وصول مفتوح

ECoMEM: Explicit Concept Memory for Memory-Dependent Robot Control

Yize Liu, Ke Wang, Mac Schwager وآخرون · 2026

A robot may lose sight of an object it must later retrieve, need to recall what a person demonstrated earlier, or track which steps of a task it has already completed. Current vision-language-action (VLA) policies often fail once the information needed for action disappears from the current observation, making memory c …

نسخة أولية وصول مفتوح

T$^2$Mem: Learning Test-Time Memory for Robotics

Yize Liu, Huang Huang, Yining Hong وآخرون · 2026

Memory-dependent robotic manipulation requires policies to use information that is no longer available in the current observation. Retaining history alone is insufficient: memory must preserve information that supports future actions. One challenge is whether a memory-free foundation model can learn to retain and use h …

نسخة أولية وصول مفتوح

CARAT: Do Materials LLMs Reason or Recite?

Jiajun Wu, Jian Yang, Zixiang Ni وآخرون · 2026

When a materials LLM answers a question about crystal structure, does it reason from the structure or copy an answer already printed in its input? Accuracy cannot tell: a structural description often prints the very field it is scored against. CARAT holds question and gold answer fixed across eight matched views, names …

نسخة أولية وصول مفتوح

DexAgent: An Agentic Human2Sim2Robot Framework for Dexterous Manipulation with Self-Evolving Tool Library

Youhui Wang, Yunzhu Li, Li Fei-Fei وآخرون · 2026

Human videos offer a scalable source of demonstrations for dexterous robot manipulation. However, existing human-to-simulation-to-robot (Human2Sim2Robot) pipelines rely on predefined procedures that struggle to accommodate diverse object properties and interactions, particularly those involving articulated and deformab …

المؤلفون المشاركون