Abstract
Conversational speech depends on dialogue context and the listener's immediately preceding behavior. We propose ReACT-TTS, a two-stage framework that uses a one-second pre-response listener facial sequence to plan the next utterance's emotion and prosody before speech realization. On a strict dyadic MELD protocol, Temporal conditioning yields higher mean macro-F1 and VAD concordance than Text-only across ten seeds, while accuracy remains essentially unchanged. Ablations show that temporal modeling performs best among the visual variants and that an explicit early-to-late difference is unnecessary; correct listener reactions also outperform cyclic mismatches on average. In a contextual-appropriateness study with 20 speech researchers, 76% of judgments prefer Temporal, 9% Text-only, and 15% report no preference. We further connect the predicted response style to a Grad-TTS backbone for end-to-end speech realization. Overall, the results support pre-response listener dynamics as complementary cues for conversational response planning. The source code is available at https://github.com/CYJ1/ReACT-TTS_public.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Chu, Y. (2026). Listen Before You Speak: Response Planning from Listener Facial Reactions for Conversational Speech Generation. https://omanscience.com/en/articles/listen-before-you-speak-response-planning-from-listener-facial-reactions-for-conversational-speech-generation
MLA 9
Chu, Yunji. "Listen Before You Speak: Response Planning from Listener Facial Reactions for Conversational Speech Generation." https://omanscience.com/en/articles/listen-before-you-speak-response-planning-from-listener-facial-reactions-for-conversational-speech-generation.
Chicago (author–date)
Chu, Yunji. 2026. "Listen Before You Speak: Response Planning from Listener Facial Reactions for Conversational Speech Generation." https://omanscience.com/en/articles/listen-before-you-speak-response-planning-from-listener-facial-reactions-for-conversational-speech-generation.
Harvard
Chu, Y. (2026) 'Listen Before You Speak: Response Planning from Listener Facial Reactions for Conversational Speech Generation', Available at: https://omanscience.com/en/articles/listen-before-you-speak-response-planning-from-listener-facial-reactions-for-conversational-speech-generation.
Vancouver
Chu Y. Listen Before You Speak: Response Planning from Listener Facial Reactions for Conversational Speech Generation. https://omanscience.com/en/articles/listen-before-you-speak-response-planning-from-listener-facial-reactions-for-conversational-speech-generation
IEEE
Y. Chu, "Listen Before You Speak: Response Planning from Listener Facial Reactions for Conversational Speech Generation," https://omanscience.com/en/articles/listen-before-you-speak-response-planning-from-listener-facial-reactions-for-conversational-speech-generation.