الباحثون

Bo Zhao

المنشورات 4

نسخة أولية وصول مفتوح

EgoFound3R: End-to-End Egocentric Hand Reconstruction in World Space with Point-Wise Interaction Attributes

Hongming Fu, Jingcheng Shi, Wenjia Wang وآخرون · 2026

Egocentric video has become a primary source of supervision for embodied models, and its value rests on recovering hand motion in world coordinates, which camera motion and hand occlusion make difficult. Existing reconstruction pipelines typically separate hand and scene estimation, leave interaction attributes to sepa …

نسخة أولية وصول مفتوح

ATI-VLA: Action-Centric Predictive Vision-Language-Action Models via Actionable Alignment Then Adaptive Injection

Yijie Zhu, Rui Shao, Jie He وآخرون · 2026

Predictive Vision-Language-Action (VLA) models aim to improve robotic manipulation via future observation or world dynamics forecasting. However, existing approaches often fail to realize this potential and underperform direct action prediction models. We argue that these limitations stem from modality misalignment bet …

نسخة أولية وصول مفتوح

FLOW: Feature-Level Optimal Warping for Generalized Remote Physiological Measurement

Bo Zhao, Junzhe Cao, Dan Guo وآخرون · 2026

Remote photoplethysmography (rPPG) enables non-contact physiological measurement but remains vulnerable to domain shifts from illumination, motion, and sensors. We propose \textbf{FLOW (Feature-Level Optimal Warping)}, an \emph{optimal transport--driven} framework for domain-generalized rPPG. FLOW integrates a \textbf{ …

نسخة أولية وصول مفتوح

Imagine-RL: Residual-Confidence-Guided Cross-Attention for World-Model-Augmented VLA Reinforcement Learning

Kejia Hu, Wentong Zhai, Bo Zhao وآخرون · 2026

Reliable action evaluation in contact-rich manipulation requires looking beyond the current observation to future visual and contact consequences. Existing noise-space reinforcement learning efficiently steers a frozen Vision-Language-Action (VLA) policy, but its critics largely ignore these consequences. We present Im …

المؤلفون المشاركون