Abstract
TTS systems with autoregressive semantic modeling have demonstrated strong zero-shot voice cloning performance and rich expressive variation, but their sequential decoding incurs substantial latency. Non-autoregressive alternatives offer much faster generation, yet often rely on more restrictive reference conditioning, such as requiring transcripts of the reference speech during inference. We present Tacit-TTS, an efficient transcript-free zero-shot voice cloning system distilled from IndexTTS2. Our model replaces autoregressive text-to-semantic decoding with masked non-autoregressive generation, introduces training-free acoustic length estimation, and accelerates the flow-matching renderer through ReFlow distillation. Across two English and two Mandarin datasets, Tacit-TTS achieves competitive zero-shot quality while generating speech over 10x faster than IndexTTS2 for utterances longer than 5 seconds. Its transcript-free conditioning further supports cross-lingual and non-lexical references. We validate this capability using references from eight other languages, infant babble, and synthetic gibberish, where transcript-dependent systems often degrade or fail due to unreliable ASR transcripts.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Chen, J., Zhang, Y., & Vinton, M. (2026). Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning. https://omanscience.com/en/articles/tacit-tts-from-autoregressive-decoding-to-masked-prediction-for-efficient-transcript-free-voice-cloning
MLA 9
Chen, Jian, et al. "Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning." https://omanscience.com/en/articles/tacit-tts-from-autoregressive-decoding-to-masked-prediction-for-efficient-transcript-free-voice-cloning.
Chicago (author–date)
Chen, Jian, You Zhang, and Mark Vinton. 2026. "Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning." https://omanscience.com/en/articles/tacit-tts-from-autoregressive-decoding-to-masked-prediction-for-efficient-transcript-free-voice-cloning.
Harvard
Chen, J., Zhang, Y. and Vinton, M. (2026) 'Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning', Available at: https://omanscience.com/en/articles/tacit-tts-from-autoregressive-decoding-to-masked-prediction-for-efficient-transcript-free-voice-cloning.
Vancouver
Chen J, Zhang Y, Vinton M. Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning. https://omanscience.com/en/articles/tacit-tts-from-autoregressive-decoding-to-masked-prediction-for-efficient-transcript-free-voice-cloning
IEEE
J. Chen, Y. Zhang, and M. Vinton, "Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning," https://omanscience.com/en/articles/tacit-tts-from-autoregressive-decoding-to-masked-prediction-for-efficient-transcript-free-voice-cloning.