الملخص
While modern text-to-speech (TTS) systems generate highly natural speech and support inline non-verbal vocalization (NVV) tags, accurate control over these events remains challenging. A key gap is the lack of established post-training methods for non-verbal control in continuous autoregressive flow-matching TTS. To this end, we present NVAlign, a direct-gradient post-training framework for NVV tag-following in this architecture. We first perform supervised fine-tuning (SFT) of TTS models and an NVV-aware automatic speech recognition (NV-ASR) model on NVV-annotated speech, then freeze the NV-ASR model to serve as the reward model for post-training. A two-step gradient surrogate enables efficient reward backpropagation through the flow-matching sampler to jointly update the autoregressive backbone and acoustic flow head. Fidelity penalties and reference-velocity regularization help preserve speaker similarity and speech quality. Results from NVV-SuperBench and human listening evaluations show that NVAlign improves tag-following accuracy over SFT and Flow-GRPO baselines. These findings demonstrate that direct reward-gradient optimization can improve non-verbal control in continuous autoregressive flow-matching TTS. Audio samples are available at https://nvalign.github.io/.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Wang, Q., Sandoval-Segura, P., Joshi, A., Jurkonis, E., & Downie, J. (2026). NVAlign: Direct-Gradient Optimization for Non-Verbal Control in Continuous Autoregressive Flow Matching Text-to-Speech. https://omanscience.com/ar/articles/nvalign-direct-gradient-optimization-for-non-verbal-control-in-continuous-autoregressive-flow-matching-text-to-speech
MLA 9
Wang, Qiaolin, et al. "NVAlign: Direct-Gradient Optimization for Non-Verbal Control in Continuous Autoregressive Flow Matching Text-to-Speech." https://omanscience.com/ar/articles/nvalign-direct-gradient-optimization-for-non-verbal-control-in-continuous-autoregressive-flow-matching-text-to-speech.
شيكاغو (المؤلف–التاريخ)
Wang, Qiaolin, Pedro Sandoval-Segura, Anunaya Joshi, Edvardas Jurkonis, and Jake Downie. 2026. "NVAlign: Direct-Gradient Optimization for Non-Verbal Control in Continuous Autoregressive Flow Matching Text-to-Speech." https://omanscience.com/ar/articles/nvalign-direct-gradient-optimization-for-non-verbal-control-in-continuous-autoregressive-flow-matching-text-to-speech.
هارفارد
Wang, Q., Sandoval-Segura, P., Joshi, A., Jurkonis, E. and Downie, J. (2026) 'NVAlign: Direct-Gradient Optimization for Non-Verbal Control in Continuous Autoregressive Flow Matching Text-to-Speech', Available at: https://omanscience.com/ar/articles/nvalign-direct-gradient-optimization-for-non-verbal-control-in-continuous-autoregressive-flow-matching-text-to-speech.
فانكوفر
Wang Q, Sandoval-Segura P, Joshi A, Jurkonis E, Downie J. NVAlign: Direct-Gradient Optimization for Non-Verbal Control in Continuous Autoregressive Flow Matching Text-to-Speech. https://omanscience.com/ar/articles/nvalign-direct-gradient-optimization-for-non-verbal-control-in-continuous-autoregressive-flow-matching-text-to-speech
IEEE
Q. Wang, P. Sandoval-Segura, A. Joshi, E. Jurkonis, and J. Downie, "NVAlign: Direct-Gradient Optimization for Non-Verbal Control in Continuous Autoregressive Flow Matching Text-to-Speech," https://omanscience.com/ar/articles/nvalign-direct-gradient-optimization-for-non-verbal-control-in-continuous-autoregressive-flow-matching-text-to-speech.