الملخص

TTS systems with autoregressive semantic modeling have demonstrated strong zero-shot voice cloning performance and rich expressive variation, but their sequential decoding incurs substantial latency. Non-autoregressive alternatives offer much faster generation, yet often rely on more restrictive reference conditioning, such as requiring transcripts of the reference speech during inference. We present Tacit-TTS, an efficient transcript-free zero-shot voice cloning system distilled from IndexTTS2. Our model replaces autoregressive text-to-semantic decoding with masked non-autoregressive generation, introduces training-free acoustic length estimation, and accelerates the flow-matching renderer through ReFlow distillation. Across two English and two Mandarin datasets, Tacit-TTS achieves competitive zero-shot quality while generating speech over 10x faster than IndexTTS2 for utterances longer than 5 seconds. Its transcript-free conditioning further supports cross-lingual and non-lexical references. We validate this capability using references from eight other languages, infant babble, and synthetic gibberish, where transcript-dependent systems often degrade or fail due to unreliable ASR transcripts.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Chen, J., Zhang, Y., & Vinton, M. (2026). Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning. https://omanscience.com/ar/articles/tacit-tts-from-autoregressive-decoding-to-masked-prediction-for-efficient-transcript-free-voice-cloning

MLA 9

Chen, Jian, et al. "Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning." https://omanscience.com/ar/articles/tacit-tts-from-autoregressive-decoding-to-masked-prediction-for-efficient-transcript-free-voice-cloning.

شيكاغو (المؤلف–التاريخ)

Chen, Jian, You Zhang, and Mark Vinton. 2026. "Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning." https://omanscience.com/ar/articles/tacit-tts-from-autoregressive-decoding-to-masked-prediction-for-efficient-transcript-free-voice-cloning.

هارفارد

Chen, J., Zhang, Y. and Vinton, M. (2026) 'Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning', Available at: https://omanscience.com/ar/articles/tacit-tts-from-autoregressive-decoding-to-masked-prediction-for-efficient-transcript-free-voice-cloning.

فانكوفر

Chen J, Zhang Y, Vinton M. Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning. https://omanscience.com/ar/articles/tacit-tts-from-autoregressive-decoding-to-masked-prediction-for-efficient-transcript-free-voice-cloning

IEEE

J. Chen, Y. Zhang, and M. Vinton, "Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning," https://omanscience.com/ar/articles/tacit-tts-from-autoregressive-decoding-to-masked-prediction-for-efficient-transcript-free-voice-cloning.