Abstract

Lightweight audio-visual speech enhancement (AVSE) models face a critical trade-off between computational efficiency and cross-modal alignment accuracy. While simple concatenation lacks relational expressiveness, dense cross-attention incurs computational overhead and is prone to unreliable cross-modal correspondence under strong acoustic interference. We propose Sparse Graph-Guided Mamba (SG-Mamba), a lightweight AVSE framework that integrates a sparse heterogeneous graph with a linear-complexity Mamba backbone. The graph explicitly models modality-specific relations through content-adaptive attention and cross-frame audio-visual connections, while Mamba captures long-range temporal context. We further introduce an audio skip connection to preserve spectral detail without sacrificing noise suppression. Evaluated on LRS3, SG-Mamba achieves competitive or superior performance against strong lightweight baselines and reaches 13.091 dB SI-SDR under noise-only condition. It also remains robust in cluttered multi-speaker conditions with a competitive cost of 3.45 G MACs (or 6.90 G FLOPs). Results on VoxCeleb2 further suggest that explicit structural priors improve robustness, generalizability, and computational efficiency in lightweight AVSE.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Tseng, G. R., Lee, H. S., Wang, H. M., & Chen, B. (2026). SG-Mamba: Sparse Graph-Guided Mamba for Audio-Visual Speech Enhancement. https://omanscience.com/en/articles/sg-mamba-sparse-graph-guided-mamba-for-audio-visual-speech-enhancement

MLA 9

Tseng, Guo-Ruei, et al. "SG-Mamba: Sparse Graph-Guided Mamba for Audio-Visual Speech Enhancement." https://omanscience.com/en/articles/sg-mamba-sparse-graph-guided-mamba-for-audio-visual-speech-enhancement.

Chicago (author–date)

Tseng, Guo-Ruei, Hung-Shin Lee, Hsin-Min Wang, and Berlin Chen. 2026. "SG-Mamba: Sparse Graph-Guided Mamba for Audio-Visual Speech Enhancement." https://omanscience.com/en/articles/sg-mamba-sparse-graph-guided-mamba-for-audio-visual-speech-enhancement.

Harvard

Tseng, G. R., Lee, H. S., Wang, H. M. and Chen, B. (2026) 'SG-Mamba: Sparse Graph-Guided Mamba for Audio-Visual Speech Enhancement', Available at: https://omanscience.com/en/articles/sg-mamba-sparse-graph-guided-mamba-for-audio-visual-speech-enhancement.

Vancouver

Tseng GR, Lee HS, Wang HM, Chen B. SG-Mamba: Sparse Graph-Guided Mamba for Audio-Visual Speech Enhancement. https://omanscience.com/en/articles/sg-mamba-sparse-graph-guided-mamba-for-audio-visual-speech-enhancement

IEEE

G. R. Tseng, H. S. Lee, H. M. Wang, and B. Chen, "SG-Mamba: Sparse Graph-Guided Mamba for Audio-Visual Speech Enhancement," https://omanscience.com/en/articles/sg-mamba-sparse-graph-guided-mamba-for-audio-visual-speech-enhancement.