الملخص

Full-duplex speech-to-speech (S2S) models enable natural conversational AI by allowing simultaneous listening and speaking. However, these models typically lack inherent user speech transcription, which is essential for applications such as conversation logging, accessibility features, and quality monitoring. In this work, we propose an efficient method to add streaming ASR capabilities to an existing duplex S2S model by introducing a lightweight ASR head in parallel to the agent text head. Our approach requires minimal additional parameters and no significant architectural changes to the base S2S model, enabling real-time user transcription while preserving full-duplex conversational capabilities including turn-taking and barge-in handling. Experimental results demonstrate that our method achieves streaming average WER of 10.21% on the HuggingFace Open ASR Leaderboard within the duplex S2S framework. Additionally, we show that the same architecture trained as a standalone streaming ASR model achieves competitive results (7.73% WER) compared to current SOTA models. We will open-source our training and inference code to facilitate further research in joint streaming ASR and S2S modeling.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Hu, K., Ferchichi, N., Casanova, E., Pasad, A., Rastorgueva, E., Chen, C., Koluguri, N. R., Zelasko, P., Peng, Y., Xu, H., Chen, Z., & Ginsburg, B. (2026). Enabling Streaming User Transcription in Full-Duplex Speech-to-Speech Models. https://omanscience.com/ar/articles/enabling-streaming-user-transcription-in-full-duplex-speech-to-speech-models

MLA 9

Hu, Ke, et al. "Enabling Streaming User Transcription in Full-Duplex Speech-to-Speech Models." https://omanscience.com/ar/articles/enabling-streaming-user-transcription-in-full-duplex-speech-to-speech-models.

شيكاغو (المؤلف–التاريخ)

Hu, Ke, Nourchene Ferchichi, Edresson Casanova, Ankita Pasad, Elena Rastorgueva, Chen Chen, Nithin Rao Koluguri, Piotr Zelasko, Yifan Peng, Hainan Xu, Zhehuai Chen, and Boris Ginsburg. 2026. "Enabling Streaming User Transcription in Full-Duplex Speech-to-Speech Models." https://omanscience.com/ar/articles/enabling-streaming-user-transcription-in-full-duplex-speech-to-speech-models.

هارفارد

Hu, K., Ferchichi, N., Casanova, E., Pasad, A., Rastorgueva, E., Chen, C., Koluguri, N. R., Zelasko, P., Peng, Y., Xu, H., Chen, Z. and Ginsburg, B. (2026) 'Enabling Streaming User Transcription in Full-Duplex Speech-to-Speech Models', Available at: https://omanscience.com/ar/articles/enabling-streaming-user-transcription-in-full-duplex-speech-to-speech-models.

فانكوفر

Hu K, Ferchichi N, Casanova E, Pasad A, Rastorgueva E, Chen C, et al. Enabling Streaming User Transcription in Full-Duplex Speech-to-Speech Models. https://omanscience.com/ar/articles/enabling-streaming-user-transcription-in-full-duplex-speech-to-speech-models

IEEE

K. Hu, N. Ferchichi, E. Casanova, A. Pasad, E. Rastorgueva, C. Chen, N. R. Koluguri, P. Zelasko, Y. Peng, H. Xu, Z. Chen, and B. Ginsburg, "Enabling Streaming User Transcription in Full-Duplex Speech-to-Speech Models," https://omanscience.com/ar/articles/enabling-streaming-user-transcription-in-full-duplex-speech-to-speech-models.