الباحثون

Bingshen Mu

المنشورات 2

نسخة أولية وصول مفتوح

From Script to Drama: An Agentic Framework for Controllable Multi-Speaker Dialogue TTS

Kangxiang Xia, Xinfa Zhu, HangRui Hu وآخرون · 2026

Multi-speaker dialogue TTS requires natural speech generation, consistent speaker identity, coherent cross-turn transitions, and fine-grained control of expressive attributes such as emotion, speaking rate, and loudness. These requirements are difficult to satisfy reliably with one-shot generation, especially in long-f …

نسخة أولية وصول مفتوح

FA-Bench: A Benchmark for Phone- and Word-Level Timestamp Accuracy in Forced Alignment and ASR on Clean and Noisy Speech

Wei Chu, Yuanzhe Dong, Ke Tan وآخرون · 2026

Forced alignment estimates the timestamps of each word, phone or character in speech given its transcript. Published comparisons normalize transcripts, split the data and match boundaries differently, so their numbers cannot be read together. We present FA-Bench, an open framework that fixes those choices once and rele …

المؤلفون المشاركون