Abstract

Conversational speech depends on dialogue context and the listener's immediately preceding behavior. We propose ReACT-TTS, a two-stage framework that uses a one-second pre-response listener facial sequence to plan the next utterance's emotion and prosody before speech realization. On a strict dyadic MELD protocol, Temporal conditioning yields higher mean macro-F1 and VAD concordance than Text-only across ten seeds, while accuracy remains essentially unchanged. Ablations show that temporal modeling performs best among the visual variants and that an explicit early-to-late difference is unnecessary; correct listener reactions also outperform cyclic mismatches on average. In a contextual-appropriateness study with 20 speech researchers, 76% of judgments prefer Temporal, 9% Text-only, and 15% report no preference. We further connect the predicted response style to a Grad-TTS backbone for end-to-end speech realization. Overall, the results support pre-response listener dynamics as complementary cues for conversational response planning. The source code is available at https://github.com/CYJ1/ReACT-TTS_public.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Chu, Y. (2026). Listen Before You Speak: Response Planning from Listener Facial Reactions for Conversational Speech Generation. https://omanscience.com/en/articles/listen-before-you-speak-response-planning-from-listener-facial-reactions-for-conversational-speech-generation

MLA 9

Chu, Yunji. "Listen Before You Speak: Response Planning from Listener Facial Reactions for Conversational Speech Generation." https://omanscience.com/en/articles/listen-before-you-speak-response-planning-from-listener-facial-reactions-for-conversational-speech-generation.

Chicago (author–date)

Chu, Yunji. 2026. "Listen Before You Speak: Response Planning from Listener Facial Reactions for Conversational Speech Generation." https://omanscience.com/en/articles/listen-before-you-speak-response-planning-from-listener-facial-reactions-for-conversational-speech-generation.

Harvard

Chu, Y. (2026) 'Listen Before You Speak: Response Planning from Listener Facial Reactions for Conversational Speech Generation', Available at: https://omanscience.com/en/articles/listen-before-you-speak-response-planning-from-listener-facial-reactions-for-conversational-speech-generation.

Vancouver

Chu Y. Listen Before You Speak: Response Planning from Listener Facial Reactions for Conversational Speech Generation. https://omanscience.com/en/articles/listen-before-you-speak-response-planning-from-listener-facial-reactions-for-conversational-speech-generation

IEEE

Y. Chu, "Listen Before You Speak: Response Planning from Listener Facial Reactions for Conversational Speech Generation," https://omanscience.com/en/articles/listen-before-you-speak-response-planning-from-listener-facial-reactions-for-conversational-speech-generation.