Abstract

Open-Vocabulary Audio-Visual Event Localization (OV-AVEL) labels each video segment with an event class, including classes that were never seen during training. The dominant pipeline uses a frozen multimodal foundation model (e.g. ImageBind) to embed the visual frame, the audio mel-spectrogram, and each candidate class name into a shared space, then computes two cosine similarities for each segment against each class: visual-text and audio-text. Existing methods then collapse this pair into a single scalar score with a fixed rule (geometric mean, weighted average) before taking the argmax. Instead, we compute complex-valued similarities and learn their fusion using a complex-valued neural network (CVNN). Each modality's standard representation becomes the real part of our pipeline, and a paired companion stream supplies the imaginary part. We use imaginary part of iHSV for visual modality and CycleGAN-translated phase spectrogram for audio modality as these companion streams. This results in two complex similarities, which are then fused. While the vision and audio encoders remain frozen, only the temporal-attention blocks and the fusion CVNN are trained. The four-stream complex architecture sets a new state of the art on both OV-AVEL benchmarks. On the open (unseen-class) split of OV-AVEBench we reach 66.5/59.1/54.1% Acc/Seg-F1/Event-F1 (+1.6/+4.1/+6.6 over the previously reported fine-tuned baseline), with consistent gains for seen classes as well. We also modify AVE dataset for this task and observe that our architecture reaches 60.7/51.9/50.4% Acc/Seg-F1/Event-F1, achieving state-of-the-art OV-AVEL results on it as well. We also propose a two-stream alternative, which also sees great improvements over the baseline.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Praveen, A., Jerripothula, K. R., Joshi, P., Dayal, A., & Sawant, N. (2026). Open-Vocabulary Audio-Visual Event Localization via Complex-Valued Fusion. https://omanscience.com/en/articles/open-vocabulary-audio-visual-event-localization-via-complex-valued-fusion

MLA 9

Praveen, Anirudh, et al. "Open-Vocabulary Audio-Visual Event Localization via Complex-Valued Fusion." https://omanscience.com/en/articles/open-vocabulary-audio-visual-event-localization-via-complex-valued-fusion.

Chicago (author–date)

Praveen, Anirudh, Koteswar Rao Jerripothula, Pratik Joshi, Aveen Dayal, and Neela Sawant. 2026. "Open-Vocabulary Audio-Visual Event Localization via Complex-Valued Fusion." https://omanscience.com/en/articles/open-vocabulary-audio-visual-event-localization-via-complex-valued-fusion.

Harvard

Praveen, A., Jerripothula, K. R., Joshi, P., Dayal, A. and Sawant, N. (2026) 'Open-Vocabulary Audio-Visual Event Localization via Complex-Valued Fusion', Available at: https://omanscience.com/en/articles/open-vocabulary-audio-visual-event-localization-via-complex-valued-fusion.

Vancouver

Praveen A, Jerripothula KR, Joshi P, Dayal A, Sawant N. Open-Vocabulary Audio-Visual Event Localization via Complex-Valued Fusion. https://omanscience.com/en/articles/open-vocabulary-audio-visual-event-localization-via-complex-valued-fusion

IEEE

A. Praveen, K. R. Jerripothula, P. Joshi, A. Dayal, and N. Sawant, "Open-Vocabulary Audio-Visual Event Localization via Complex-Valued Fusion," https://omanscience.com/en/articles/open-vocabulary-audio-visual-event-localization-via-complex-valued-fusion.