الباحثون

Hang Zhao

المنشورات 11

نسخة أولية وصول مفتوح

Supervising Sound Localization by In-the-wild Egomotion

Anna Min, Ziyang Chen, Hang Zhao وآخرون · 2026

We present a method for learning binaural sound localization using egomotion as a supervisory signal. Over the course of a video, the cameras direction to a sound source will change as the camera moves. We train an audio model to predict sound directions that are consistent with visual estimates of camera motion, which …

نسخة أولية وصول مفتوح

Rethinking Causal Action Tokenization with Conditional Annealing in Flow Matching

Chenyu Zhang, Yuhang Cao, Daru Du وآخرون · 2026

Autoregressive Vision-Language-Action (VLA) models offer a scalable path to robot learning, yet existing action tokenizers treat tokenization as a compression problem, producing representations that are semantically misaligned with the autoregressive backbone. We propose CATok, a causal action tokenizer that reframes t …

نسخة أولية وصول مفتوح

World SLAM Model: Joint World Modeling for SLAM and Navigation

Minghui Qin, Yijun Yuan, Weicheng Zheng وآخرون · 2026

We introduce World SLAM Model (WSM), a unified framework that brings the SLAM paradigm directly into downstream navigation. Rather than treating SLAM merely as an upstream module that provides poses, maps or tokens, WSM adopts its core mechanisms, including incremental state updates with persistent memory and backend r …

نسخة أولية وصول مفتوح

TactileStep: Sole Tactile Learning for Regulating Foot-Terrain Interaction in Humanoid Locomotion

Zizhuo Wang, Ming-Ju Lee, Shaoting Zhu وآخرون · 2026

Humanoid parkour policies can traverse various terrains, but task completion may mask challenges of harsh landings, edge contacts, and unstable stance contacts. Humans naturally regulate foot-terrain interaction through tactile feedback, modulating contact compliance according to terrain stiffness. This highlights a ke …

نسخة أولية وصول مفتوح

Representation World Model: Learning States, Transition and Executable Plans in Representation

Yijun Yuan, Weicheng Zheng, Weibang Wang وآخرون · 2026

We propose the Representation World Model (RWM), which learns states, transitions, and executable plans directly in representation space. Unlike existing world models that typically learn latent representations together with explicit dynamics models and perform planning through search, optimization, or policy-based pre …

نسخة أولية وصول مفتوح

CereVLA: Cerebellum-Inspired Consequence-Aware Residual Governance for Efficient Vision-Language-Action Execution

Shuai Zeng, Yuxuan Liang, Hangmiao Hu وآخرون · 2026

Action-chunked vision-language-action (VLA) policies improve inference efficiency, but limited feedback within committed action chunks can lead to accumulated execution errors. Residual adaptation can correct such deviations without retraining the VLA; however, existing corrections are typically optimized for reference …

نسخة أولية وصول مفتوح

MSK-Bench: Benchmarking Full-Body Musculoskeletal Motor Control Across Tasks, Control Paradigms, and Physiological Metrics

Mengtao Ou, Zongzheng Zhang, Zhenghao Xiao وآخرون · 2026

Musculoskeletal (MSK) humanoids provide a physiologically grounded embodiment for studying full-body motor control, but their high-dimensional muscle actuation, delayed activation dynamics, and redundant muscle--tendon structures make learning substantially harder than torque-driven humanoid control. Existing MSK bench …

نسخة أولية وصول مفتوح

SkelWAM: A Skeleton-Guided World-Action Model for Zero-Shot Cross-Embodiment Manipulation

Pengjun Niu, Yujia Xie, Rui Peng وآخرون · 2026

Reusing manipulation experience across robot embodiments is important for scaling robot learning and reducing repeated task-specific data collection. However, changes in embodiment alter visual appearance, action dimensionality and semantics, and the whole-body configurations that can realize the same tool pose. We pre …

المؤلفون المشاركون