الباحثون

Chunhua Shen

المنشورات 6

نسخة أولية وصول مفتوح

PerturBot: Breaking Shortcut Priors in Vision-Language-Action Models with Perturbative Training

Mingyu Liu, Chonghao Sima, Tianjian Feng وآخرون · 2026

A vision--language--action (VLA) policy can complete complex tasks while ignoring the evidence that should determine its actions. An object held near the wrist camera can displace the instructed target. Language and action show the same pattern: a familiar noun can trigger the operation it was paired with in training e …

نسخة أولية وصول مفتوح

GeoVerse: World-Consistent Novel View Synthesis in Geometric Latent Space

Kerui Ren, Tao Lu, Linning Xu وآخرون · 2026

Novel view synthesis from sparse images must reconcile faithful reconstruction of observed regions with plausible completion of unseen content, while maintaining world consistency across viewpoints. Existing geometry-based methods preserve observed scene structure but often struggle to complete unseen regions, whereas …

نسخة أولية وصول مفتوح

InfiniHand: Streaming World-Space Hand Motion Estimation from Egocentric Video

Kerui Ren, Kaiwen Song, Weiguang Zhao وآخرون · 2026

World-space hand motion estimation from egocentric video requires recovering 3D articulated hand geometry while tracking camera egomotion. Existing approaches heavily rely on cascading independent hand pose estimators and SLAM systems, resulting in error accumulation, complex pipelines, and severe computational overhea …

نسخة أولية وصول مفتوح

InternW0-$Δ$: A World Action Model Bridging Predictive Dynamics and Actions with 20K+ Hours of Open Data

Xingyu Miao, Zizun Li, Baole Fang وآخرون · 2026

World Action Models (WAMs) jointly model visual dynamics and action generation for generalist robot manipulation. A central challenge is to integrate priors from large-scale pretrained models---including visual dynamics, scene semantics, geometry, and motion---into a unified framework for robot action generation. We in …

نسخة أولية وصول مفتوح

InternW0: A Foundational Physical World Model for Efficient Real-World Interactions

Jisong Cai, Yao Mu, Ganlin Yang وآخرون · 2026

Physical intelligence requires more than predicting how the world may evolve: predictions must remain actionable as the world continues to change. We introduce InternW0, the first instantiation of the InternW physical world model series from Shanghai AI Laboratory, built around omnimodal interfaces, asynchronous multi- …

نسخة أولية وصول مفتوح

Metric-Bench: Exploring In-context Spatial Metric Reasoning in VLMs for Indoor Scenes

Yuling Xi, Haokai Zhang, Muzhi Zhu وآخرون · 2026

Metric reasoning is a critical and challenging task for Vision Language Models (VLMs), playing a pivotal role in embodied AI tasks such as robotic manipulation and autonomous navigation. However, current spatial reasoning remains bottlenecked by rigid pixel-level supervision; such localized optimization often compromis …

المؤلفون المشاركون