Abstract

Video-to-speech synthesis aims to generate natural-sounding speech from silent talking-face videos while ensuring phonetic accuracy. A fundamental challenge in this task is the inherent one-to-many mapping problem, where visual dynamics often lack sufficient information to uniquely determine the corresponding utterance. To address this, we propose Watch Your Speech (WYS), a video-to-speech synthesis framework that incorporates textual conditioning as an explicit linguistic cue to mitigate visual ambiguity. Our framework features an attention-based embedding fusion module that synergistically integrates textual context with video sequences, coupled with a conditional flow matching objective for high-fidelity speech generation. Extensive experiments on the LRS2 and LRS3 datasets demonstrate that WYS achieves superior performance, establishing new state-of-the-art results in audio-visual synchronization (LSE-C/D) while maintaining highly competitive textual accuracy (WER). Subjective evaluations further confirm that our model generates speech with near-human naturalness, validating the effectiveness of textual conditioning in content-controlled video-to-speech synthesis. Project page: https://github.com/gunwoo5034/Watch-your-Speech

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Lee, G., Oh, Y., & Han, Y. (2026). Watch Your Speech: Text-aware Video-to-Speech Synthesis with Textual Conditioning. https://omanscience.com/en/articles/watch-your-speech-text-aware-video-to-speech-synthesis-with-textual-conditioning

MLA 9

Lee, Gunwoo, et al. "Watch Your Speech: Text-aware Video-to-Speech Synthesis with Textual Conditioning." https://omanscience.com/en/articles/watch-your-speech-text-aware-video-to-speech-synthesis-with-textual-conditioning.

Chicago (author–date)

Lee, Gunwoo, Yoori Oh, and Yoseob Han. 2026. "Watch Your Speech: Text-aware Video-to-Speech Synthesis with Textual Conditioning." https://omanscience.com/en/articles/watch-your-speech-text-aware-video-to-speech-synthesis-with-textual-conditioning.

Harvard

Lee, G., Oh, Y. and Han, Y. (2026) 'Watch Your Speech: Text-aware Video-to-Speech Synthesis with Textual Conditioning', Available at: https://omanscience.com/en/articles/watch-your-speech-text-aware-video-to-speech-synthesis-with-textual-conditioning.

Vancouver

Lee G, Oh Y, Han Y. Watch Your Speech: Text-aware Video-to-Speech Synthesis with Textual Conditioning. https://omanscience.com/en/articles/watch-your-speech-text-aware-video-to-speech-synthesis-with-textual-conditioning

IEEE

G. Lee, Y. Oh, and Y. Han, "Watch Your Speech: Text-aware Video-to-Speech Synthesis with Textual Conditioning," https://omanscience.com/en/articles/watch-your-speech-text-aware-video-to-speech-synthesis-with-textual-conditioning.