Abstract
Lightweight audio-visual speech enhancement (AVSE) models face a critical trade-off between computational efficiency and cross-modal alignment accuracy. While simple concatenation lacks relational expressiveness, dense cross-attention incurs computational overhead and is prone to unreliable cross-modal correspondence under strong acoustic interference. We propose Sparse Graph-Guided Mamba (SG-Mamba), a lightweight AVSE framework that integrates a sparse heterogeneous graph with a linear-complexity Mamba backbone. The graph explicitly models modality-specific relations through content-adaptive attention and cross-frame audio-visual connections, while Mamba captures long-range temporal context. We further introduce an audio skip connection to preserve spectral detail without sacrificing noise suppression. Evaluated on LRS3, SG-Mamba achieves competitive or superior performance against strong lightweight baselines and reaches 13.091 dB SI-SDR under noise-only condition. It also remains robust in cluttered multi-speaker conditions with a competitive cost of 3.45 G MACs (or 6.90 G FLOPs). Results on VoxCeleb2 further suggest that explicit structural priors improve robustness, generalizability, and computational efficiency in lightweight AVSE.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Tseng, G. R., Lee, H. S., Wang, H. M., & Chen, B. (2026). SG-Mamba: Sparse Graph-Guided Mamba for Audio-Visual Speech Enhancement. https://omanscience.com/en/articles/sg-mamba-sparse-graph-guided-mamba-for-audio-visual-speech-enhancement
MLA 9
Tseng, Guo-Ruei, et al. "SG-Mamba: Sparse Graph-Guided Mamba for Audio-Visual Speech Enhancement." https://omanscience.com/en/articles/sg-mamba-sparse-graph-guided-mamba-for-audio-visual-speech-enhancement.
Chicago (author–date)
Tseng, Guo-Ruei, Hung-Shin Lee, Hsin-Min Wang, and Berlin Chen. 2026. "SG-Mamba: Sparse Graph-Guided Mamba for Audio-Visual Speech Enhancement." https://omanscience.com/en/articles/sg-mamba-sparse-graph-guided-mamba-for-audio-visual-speech-enhancement.
Harvard
Tseng, G. R., Lee, H. S., Wang, H. M. and Chen, B. (2026) 'SG-Mamba: Sparse Graph-Guided Mamba for Audio-Visual Speech Enhancement', Available at: https://omanscience.com/en/articles/sg-mamba-sparse-graph-guided-mamba-for-audio-visual-speech-enhancement.
Vancouver
Tseng GR, Lee HS, Wang HM, Chen B. SG-Mamba: Sparse Graph-Guided Mamba for Audio-Visual Speech Enhancement. https://omanscience.com/en/articles/sg-mamba-sparse-graph-guided-mamba-for-audio-visual-speech-enhancement
IEEE
G. R. Tseng, H. S. Lee, H. M. Wang, and B. Chen, "SG-Mamba: Sparse Graph-Guided Mamba for Audio-Visual Speech Enhancement," https://omanscience.com/en/articles/sg-mamba-sparse-graph-guided-mamba-for-audio-visual-speech-enhancement.