الملخص

We present JoyAI-Voice~2.0, an end-to-end anthropomorphic speech generation model built upon a fully continuous, dual-encoder architecture. Raw speech is encoded into continuous latents and partitioned into patches. Each patch is decomposed by a semantic-acoustic dual encoder into a semantically purified representation and an acoustic representation, which are fused and jointly fed to a causal autoregressive Transformer for planning. The Transformer predicts the conditioning for the next patch, and a local diffusion Transformer renders its full latents for 48\,kHz synthesis. The model is trained with a joint flow-matching and stop-prediction objective, followed by supervised fine-tuning and reinforcement learning with DiffusionNFT to improve model performance. It achieves the lowest average word error rate of 2.51\% on Seed-TTS, a 14.9\% relative reduction over the strongest baseline, and state-of-the-art attribute fidelity on InstructTTSEval in both Chinese and English, leading on 5 of 10 perceptual dimensions with the highest overall score of 0.893 on MDVD-Eval.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Chen, Y., Dong, B., Huang, Y., Li, H., Li, J., Liang, X., Ni, H., Wang, W., Wang, Y., Xiao, Z., Deng, W., Duan, N., Gu, Y., Guan, W., Han, W., Li, Y., Liu, Y., Ye, J., Yu, F., & Zhu, L. (2026). JoyAI-Voice 2.0: A Full-Continuous Autoregressive Speech Generation Model with Semantic-Acoustic Joint Representation. https://omanscience.com/ar/articles/joyai-voice-2-0-a-full-continuous-autoregressive-speech-generation-model-with-semantic-acoustic-joint-representation

MLA 9

Chen, Yafeng, et al. "JoyAI-Voice 2.0: A Full-Continuous Autoregressive Speech Generation Model with Semantic-Acoustic Joint Representation." https://omanscience.com/ar/articles/joyai-voice-2-0-a-full-continuous-autoregressive-speech-generation-model-with-semantic-acoustic-joint-representation.

شيكاغو (المؤلف–التاريخ)

Chen, Yafeng, Boya Dong, Yankun Huang, Hao Li, Jingdong Li, Xiangyu Liang, Hao Ni, Wenchao Wang, Yuxuan Wang, Zhangyu Xiao, Wei Deng, Nan Duan, Yu Gu, Wenhao Guan, Weisheng Han, Yabin Li, Yuan Liu, Jiaxin Ye, Fan Yu, and Lin Zhu. 2026. "JoyAI-Voice 2.0: A Full-Continuous Autoregressive Speech Generation Model with Semantic-Acoustic Joint Representation." https://omanscience.com/ar/articles/joyai-voice-2-0-a-full-continuous-autoregressive-speech-generation-model-with-semantic-acoustic-joint-representation.

هارفارد

Chen, Y., Dong, B., Huang, Y., Li, H., Li, J., Liang, X., Ni, H., Wang, W., Wang, Y., Xiao, Z., Deng, W., Duan, N., Gu, Y., Guan, W., Han, W., Li, Y., Liu, Y., Ye, J., Yu, F. and Zhu, L. (2026) 'JoyAI-Voice 2.0: A Full-Continuous Autoregressive Speech Generation Model with Semantic-Acoustic Joint Representation', Available at: https://omanscience.com/ar/articles/joyai-voice-2-0-a-full-continuous-autoregressive-speech-generation-model-with-semantic-acoustic-joint-representation.

فانكوفر

Chen Y, Dong B, Huang Y, Li H, Li J, Liang X, et al. JoyAI-Voice 2.0: A Full-Continuous Autoregressive Speech Generation Model with Semantic-Acoustic Joint Representation. https://omanscience.com/ar/articles/joyai-voice-2-0-a-full-continuous-autoregressive-speech-generation-model-with-semantic-acoustic-joint-representation

IEEE

Y. Chen, B. Dong, Y. Huang, H. Li, J. Li, X. Liang, H. Ni, W. Wang, Y. Wang, Z. Xiao, W. Deng, N. Duan, Y. Gu, W. Guan, W. Han, Y. Li, Y. Liu, J. Ye, F. Yu, and L. Zhu, "JoyAI-Voice 2.0: A Full-Continuous Autoregressive Speech Generation Model with Semantic-Acoustic Joint Representation," https://omanscience.com/ar/articles/joyai-voice-2-0-a-full-continuous-autoregressive-speech-generation-model-with-semantic-acoustic-joint-representation.