الباحثون

Cheng Peng

المنشورات 3

نسخة أولية وصول مفتوح

Less Context, Better Geometry: Masked Geometric Encoder for Robust 3D Foundation Models

Zhimin Shao, Xijun Liu, Zhaoliang Zhang وآخرون · 2026

Recent progress in 3D foundation models has enabled rapid 3D reconstruction and camera calibration by leveraging learned 3D priors from vast amount of spatial data. However, the all-to-all global attention design leads to quadratic complexity and limits long-sequence inference; unconstrained cross-view interactions als …

نسخة أولية وصول مفتوح

DIDO: Distilling Interaction-Centric Dynamics into One-Step Denoising for World Action Models

Jing Lyu, Shuanghao Bai, Runze Xiao وآخرون · 2026

World Action Models (WAMs) use video generation models to predict future visual dynamics for robotic manipulation, but iterative denoising introduces additional latency for closed-loop control. We empirically find that visual content converges at different rates during denoising. Static background structure forms early …

المؤلفون المشاركون