الملخص
Generating full-body co-speech motion for humanoid robots requires coordinating speech prosody, linguistic content, and embodiment-specific motion. To this end, we present ECHO-G, a framework jointly conditioned on speech audio and timed transcripts. Its Speech-Grounded Diffusion Transformer (SGDiT) combines frame-aligned acoustic features with token-level linguistic context, preserving their distinct granularities. Trained with rectified flow matching, it models one-to-many utterance-motion relationships directly in robot space. To support training and evaluation, we introduce a BEAT2-derived audio-text-robot dataset and a benchmark covering co-speech characteristics, robot-motion quality, and runtime efficiency. Comparative evaluation supports direct robot-space generation over the evaluated human-motion generation and retargeting pipelines, while modality ablations highlight the benefits of joint audio-text conditioning. We further demonstrate deployment on a physical humanoid robot. A complementary video-rating study also favors joint conditioning over the alternatives. The dataset and training, inference, and evaluation code are available through our project page.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Li, Y., Gao, P., Wang, M., Shen, S., Yang, S., & Xu, H. (2026). ECHO-G: Embodied Co-speech Humanoid mOtion Generation. https://omanscience.com/ar/articles/echo-g-embodied-co-speech-humanoid-motion-generation
MLA 9
Li, Yizhao, et al. "ECHO-G: Embodied Co-speech Humanoid mOtion Generation." https://omanscience.com/ar/articles/echo-g-embodied-co-speech-humanoid-motion-generation.
شيكاغو (المؤلف–التاريخ)
Li, Yizhao, Pusen Gao, Ming Wang, Shaojie Shen, Shuo Yang, and Hao Xu. 2026. "ECHO-G: Embodied Co-speech Humanoid mOtion Generation." https://omanscience.com/ar/articles/echo-g-embodied-co-speech-humanoid-motion-generation.
هارفارد
Li, Y., Gao, P., Wang, M., Shen, S., Yang, S. and Xu, H. (2026) 'ECHO-G: Embodied Co-speech Humanoid mOtion Generation', Available at: https://omanscience.com/ar/articles/echo-g-embodied-co-speech-humanoid-motion-generation.
فانكوفر
Li Y, Gao P, Wang M, Shen S, Yang S, Xu H. ECHO-G: Embodied Co-speech Humanoid mOtion Generation. https://omanscience.com/ar/articles/echo-g-embodied-co-speech-humanoid-motion-generation
IEEE
Y. Li, P. Gao, M. Wang, S. Shen, S. Yang, and H. Xu, "ECHO-G: Embodied Co-speech Humanoid mOtion Generation," https://omanscience.com/ar/articles/echo-g-embodied-co-speech-humanoid-motion-generation.