Abstract

Multi-speaker dialogue TTS requires natural speech generation, consistent speaker identity, coherent cross-turn transitions, and fine-grained control of expressive attributes such as emotion, speaking rate, and loudness. These requirements are difficult to satisfy reliably with one-shot generation, especially in long-form dialogue. We propose a controllable multi-speaker dialogue TTS framework that formulates synthesis as critique-driven iterative refinement. Its speech backbone, ControlEdit-TTS, unifies instruction-following synthesis and natural-language-guided attribute editing, enabling correction of expressive errors without full regeneration. The framework further performs hierarchical utterance-level and scene-level critique, routing detected issues to editing, resynthesis, or timing adjustment. Experiments on a bilingual Chinese--English dialogue benchmark show improved utterance-level instruction following, better dialogue-level preference than direct dialogue models and agentic baselines, and more effective refinement than regeneration-only alternatives while preserving speaker identity. Ablations further confirm the benefits of scene-level critique and edit-based correction.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Xia, K., Zhu, X., Hu, H., Huang, K., Tian, W., Jiang, Z., Mu, B., Hu, J., He, T., Xie, L., & Xu, J. (2026). From Script to Drama: An Agentic Framework for Controllable Multi-Speaker Dialogue TTS. https://omanscience.com/en/articles/from-script-to-drama-an-agentic-framework-for-controllable-multi-speaker-dialogue-tts

MLA 9

Xia, Kangxiang, et al. "From Script to Drama: An Agentic Framework for Controllable Multi-Speaker Dialogue TTS." https://omanscience.com/en/articles/from-script-to-drama-an-agentic-framework-for-controllable-multi-speaker-dialogue-tts.

Chicago (author–date)

Xia, Kangxiang, Xinfa Zhu, HangRui Hu, Kexin Huang, Wenjie Tian, Ziyue Jiang, Bingshen Mu, Jingbin Hu, Ting He, Lei Xie, and Jin Xu. 2026. "From Script to Drama: An Agentic Framework for Controllable Multi-Speaker Dialogue TTS." https://omanscience.com/en/articles/from-script-to-drama-an-agentic-framework-for-controllable-multi-speaker-dialogue-tts.

Harvard

Xia, K., Zhu, X., Hu, H., Huang, K., Tian, W., Jiang, Z., Mu, B., Hu, J., He, T., Xie, L. and Xu, J. (2026) 'From Script to Drama: An Agentic Framework for Controllable Multi-Speaker Dialogue TTS', Available at: https://omanscience.com/en/articles/from-script-to-drama-an-agentic-framework-for-controllable-multi-speaker-dialogue-tts.

Vancouver

Xia K, Zhu X, Hu H, Huang K, Tian W, Jiang Z, et al. From Script to Drama: An Agentic Framework for Controllable Multi-Speaker Dialogue TTS. https://omanscience.com/en/articles/from-script-to-drama-an-agentic-framework-for-controllable-multi-speaker-dialogue-tts

IEEE

K. Xia, X. Zhu, H. Hu, K. Huang, W. Tian, Z. Jiang, B. Mu, J. Hu, T. He, L. Xie, and J. Xu, "From Script to Drama: An Agentic Framework for Controllable Multi-Speaker Dialogue TTS," https://omanscience.com/en/articles/from-script-to-drama-an-agentic-framework-for-controllable-multi-speaker-dialogue-tts.