الملخص

Behavioral foundation models (BFMs) have recently shown that a single humanoid policy can support diverse whole-body control, but extending such generality to physical interaction remains challenging. We introduce I-BFM, to our knowledge the first BFM for humanoid-object interaction. Rather than relying on task-specific policies or reference tracking, I-BFM learns a shared representation of the coupled dynamics among the humanoid, objects, and their contacts using forward-backward representations and unsupervised reinforcement learning. Given a downstream task reward, the same policy can be directly conditioned on a latent command to execute closed-loop interaction without task-specific policy optimization. To improve interaction control over different time scales, we further train the policy with both short-horizon interaction targets and longer-horizon goal targets. A single I-BFM policy performs carrying, pushing, and kicking, while also supporting goal reaching, motion tracking, stylistic control, and long-horizon task chaining. More importantly, it remains effective after large deviations from nominal execution: on Carry, I-BFM achieves 94.3% nominal success and retains 89.3% success after robot falls, compared with 1.3% for a planning-based baseline. Real-world experiments on a Unitree G1 further demonstrate diverse loco-manipulation behaviors, rapid recovery from interaction failures and external disturbances, and task chaining without task-specific retraining.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Han, Z., Li, Y., Sun, J., Dong, F., Shen, Y., Ye, L., Jing, Z., Zhang, Y., Zhang, Y., Wang, X., & Zhao, H. (2026). I-BFM: Reward-Conditioned Robust Humanoid Interaction via Unsupervised Reinforcement Learning. https://omanscience.com/ar/articles/i-bfm-reward-conditioned-robust-humanoid-interaction-via-unsupervised-reinforcement-learning

MLA 9

Han, Ziqi, et al. "I-BFM: Reward-Conditioned Robust Humanoid Interaction via Unsupervised Reinforcement Learning." https://omanscience.com/ar/articles/i-bfm-reward-conditioned-robust-humanoid-interaction-via-unsupervised-reinforcement-learning.

شيكاغو (المؤلف–التاريخ)

Han, Ziqi, Yitang Li, Junhan Sun, Fanrong Dong, Yaojie Shen, Lei Ye, Zetong Jing, Yongqi Zhang, Yiming Zhang, Xue Wang, and Hao Zhao. 2026. "I-BFM: Reward-Conditioned Robust Humanoid Interaction via Unsupervised Reinforcement Learning." https://omanscience.com/ar/articles/i-bfm-reward-conditioned-robust-humanoid-interaction-via-unsupervised-reinforcement-learning.

هارفارد

Han, Z., Li, Y., Sun, J., Dong, F., Shen, Y., Ye, L., Jing, Z., Zhang, Y., Zhang, Y., Wang, X. and Zhao, H. (2026) 'I-BFM: Reward-Conditioned Robust Humanoid Interaction via Unsupervised Reinforcement Learning', Available at: https://omanscience.com/ar/articles/i-bfm-reward-conditioned-robust-humanoid-interaction-via-unsupervised-reinforcement-learning.

فانكوفر

Han Z, Li Y, Sun J, Dong F, Shen Y, Ye L, et al. I-BFM: Reward-Conditioned Robust Humanoid Interaction via Unsupervised Reinforcement Learning. https://omanscience.com/ar/articles/i-bfm-reward-conditioned-robust-humanoid-interaction-via-unsupervised-reinforcement-learning

IEEE

Z. Han, Y. Li, J. Sun, F. Dong, Y. Shen, L. Ye, Z. Jing, Y. Zhang, Y. Zhang, X. Wang, and H. Zhao, "I-BFM: Reward-Conditioned Robust Humanoid Interaction via Unsupervised Reinforcement Learning," https://omanscience.com/ar/articles/i-bfm-reward-conditioned-robust-humanoid-interaction-via-unsupervised-reinforcement-learning.