Abstract

Head-pose variation introduces substantial appearance transformations in visual speech recognition (VSR), making pose-aware feature modulation desirable. However, performance degradation and unwanted feature interactions may result from using numerous Feature-wise Linear Modulation (FiLM) circuits with fixed modulation intensity. We propose a Pose Adaptive Dynamic FiLM framework with a Dynamic Residual FiLM (DR-FiLM) modulator that predicts input-dependent weights to adaptively control the strength of pose-conditioned modulation. Experiments on LRS2 and LRS3 demonstrate that unweighted multi-pathway modulation substantially degrades phoneme recognition, increasing PER to 20.33% and 29.42%, respectively, compared with 16.20% and 20.96% for the single ResFiLM configuration. In contrast, the proposed DR-FiLM with dynamic Deep-Res weighting reduces PER to 15.74% on LRS2 and 23.91% on LRS3, substantially mitigating the adverse effects of unweighted modulation. The analysis of the learned weights further reveals a consistent tendency to assign greater weight to the deeper FiLM pathway as head-pose variation increases. These results show that merging pose-conditioned FiLM circuits is more efficient when the modulation strength is dynamically controlled.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Teng, M. K. K., Zhang, H., & Saitoh, T. (2026). Pose Adaptive Dynamic FiLM Modulation for Visual Speech Recognition. https://omanscience.com/en/articles/pose-adaptive-dynamic-film-modulation-for-visual-speech-recognition

MLA 9

Teng, Matthew Kit Khinn, et al. "Pose Adaptive Dynamic FiLM Modulation for Visual Speech Recognition." https://omanscience.com/en/articles/pose-adaptive-dynamic-film-modulation-for-visual-speech-recognition.

Chicago (author–date)

Teng, Matthew Kit Khinn, Haibo Zhang, and Takeshi Saitoh. 2026. "Pose Adaptive Dynamic FiLM Modulation for Visual Speech Recognition." https://omanscience.com/en/articles/pose-adaptive-dynamic-film-modulation-for-visual-speech-recognition.

Harvard

Teng, M. K. K., Zhang, H. and Saitoh, T. (2026) 'Pose Adaptive Dynamic FiLM Modulation for Visual Speech Recognition', Available at: https://omanscience.com/en/articles/pose-adaptive-dynamic-film-modulation-for-visual-speech-recognition.

Vancouver

Teng MKK, Zhang H, Saitoh T. Pose Adaptive Dynamic FiLM Modulation for Visual Speech Recognition. https://omanscience.com/en/articles/pose-adaptive-dynamic-film-modulation-for-visual-speech-recognition

IEEE

M. K. K. Teng, H. Zhang, and T. Saitoh, "Pose Adaptive Dynamic FiLM Modulation for Visual Speech Recognition," https://omanscience.com/en/articles/pose-adaptive-dynamic-film-modulation-for-visual-speech-recognition.