Authors

Jiajun Wu

Publications 13

Preprint Open access

BrickBench: Evaluating Agentic Brick Design

We propose BrickBench, a benchmark for agentic text-conditioned LEGO-set design. Given a prompt, an agent is tasked with producing an assembly that not only satisfies semantic and design criteria, but that can also be physically built. To do so, it must select parts from a discrete library and reason jointly about loca …

Preprint Open access

OpenWAM: An Open Framework for Composable World-Action Models

Heng Yu, David D. Yuan, Juze Zhang et al. · 2026

World-action models (WAMs) couple future prediction with robot control, yet existing systems often vary the video backbone, interaction structure, supervision, and inference procedure simultaneously, making their design choices difficult to compare. We introduce OPENWAM, an open world-action modeling framework built ar …

Preprint Open access

Robot Learning with Visual Predicted Force

Haonan Chen, Feiyang Wu, Yuxiang Ma et al. · 2026

Force-aware manipulation typically relies on specialized force or tactile sensors. We show that force-aware manipulation can instead be achieved through visual force prediction from the deformation of a compliant Fin Ray gripper. Our approach trains two models. First, we train a visual force estimator on calibration da …

Preprint Open access

T$^2$Mem: Learning Test-Time Memory for Robotics

Yize Liu, Huang Huang, Yining Hong et al. · 2026

Memory-dependent robotic manipulation requires policies to use information that is no longer available in the current observation. Retaining history alone is insufficient: memory must preserve information that supports future actions. One challenge is whether a memory-free foundation model can learn to retain and use h …

Preprint Open access

CARAT: Do Materials LLMs Reason or Recite?

Jiajun Wu, Jian Yang, Zixiang Ni et al. · 2026

When a materials LLM answers a question about crystal structure, does it reason from the structure or copy an answer already printed in its input? Accuracy cannot tell: a structural description often prints the very field it is scored against. CARAT holds question and gold answer fixed across eight matched views, names …

Preprint Open access

DexAgent: An Agentic Human2Sim2Robot Framework for Dexterous Manipulation with Self-Evolving Tool Library

Youhui Wang, Yunzhu Li, Li Fei-Fei et al. · 2026

Human videos offer a scalable source of demonstrations for dexterous robot manipulation. However, existing human-to-simulation-to-robot (Human2Sim2Robot) pipelines rely on predefined procedures that struggle to accommodate diverse object properties and interactions, particularly those involving articulated and deformab …

Co-authors