Abstract
Emotion-conditioned text-to-speech (TTS) models may fail to express the requested emotion reliably, and improving controllability by additional training is costly in both computation and emotion-labeled speech training data. We therefore study vector steering, a training-free approach that modifies the internal representations of a frozen model. CoCoEmo, a conventional vector steering method for emotion TTS, treats each emotion vector as an indivisible direction controlled by a single global strength, limiting adherence to the requested emotion. In this work, we first discover that an emotion vector can be decomposed into a shared component that moves speech away from neutral expression and a residual component that directs generation toward the requested emotion. Building on this finding, we propose Emotion Residual-Enhanced Steering for TTS (EmoRES), a novel method that controls the two components without retraining the backbone. On IEMOCAP, EmoRES outperforms CoCoEmo across all four objective emotion metrics on the IndexTTS-2 and CosyVoice2 backbones. Rank correlation improves by 26.13 and 12.97 percentage points, corresponding to relative gains of 118.8% and 33.1%, while emotion hit rate improves by 12.95 and 6.92 points, corresponding to relative gains of 20.1% and 9.8%. Human evaluation further shows a relative improvement up to 35.0% in the rate at which listeners correctly identified the dominant requested emotion and up to a 17.3% improvement in fidelity, while listeners prefer EmoRES for naturalness in up to 63.8% of pairwise comparisons. Component ablations further demonstrate that effective control benefits from preserving the shared component while strengthening the residual of the emotion steering vectors.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Huang, K. P., Liu, H., Peng, P., Wu, H., Ni, Z., Lee, H. Y., Lee, J., & Chachra, N. (2026). EmoRES-TTS: Residual-Enhanced Vector Steering for Emotional Speech Generation. https://omanscience.com/en/articles/emores-tts-residual-enhanced-vector-steering-for-emotional-speech-generation
MLA 9
Huang, Kuan-Po, et al. "EmoRES-TTS: Residual-Enhanced Vector Steering for Emotional Speech Generation." https://omanscience.com/en/articles/emores-tts-residual-enhanced-vector-steering-for-emotional-speech-generation.
Chicago (author–date)
Huang, Kuan-Po, Haohe Liu, Puyuan Peng, Haibin Wu, Zhaoheng Ni, Hung-yi Lee, Jinwon Lee, and Neha Chachra. 2026. "EmoRES-TTS: Residual-Enhanced Vector Steering for Emotional Speech Generation." https://omanscience.com/en/articles/emores-tts-residual-enhanced-vector-steering-for-emotional-speech-generation.
Harvard
Huang, K. P., Liu, H., Peng, P., Wu, H., Ni, Z., Lee, H. Y., Lee, J. and Chachra, N. (2026) 'EmoRES-TTS: Residual-Enhanced Vector Steering for Emotional Speech Generation', Available at: https://omanscience.com/en/articles/emores-tts-residual-enhanced-vector-steering-for-emotional-speech-generation.
Vancouver
Huang KP, Liu H, Peng P, Wu H, Ni Z, Lee HY, et al. EmoRES-TTS: Residual-Enhanced Vector Steering for Emotional Speech Generation. https://omanscience.com/en/articles/emores-tts-residual-enhanced-vector-steering-for-emotional-speech-generation
IEEE
K. P. Huang, H. Liu, P. Peng, H. Wu, Z. Ni, H. Y. Lee, J. Lee, and N. Chachra, "EmoRES-TTS: Residual-Enhanced Vector Steering for Emotional Speech Generation," https://omanscience.com/en/articles/emores-tts-residual-enhanced-vector-steering-for-emotional-speech-generation.