Abstract

Large audio-language models (LALMs) are sensitive to input perturbations, such as noise, waveform corruption, and adversarial injections. We propose AnchorPrompt, an efficient adaptation method that keeps the model frozen and learns a single block of prompt vectors inserted at the decoder input, between the audio and question embeddings. We train these vectors through self-distillation over diverse audio and text perturbations. To improve answer consistency and mitigate hallucination, we use the model's prediction on the clean recording as the target for answerable inputs, and assign a refusal target when the audio lacks sufficient evidence to answer. Furthermore, AnchorPrompt is perturbation-agnostic at inference, requiring no prior detection of perturbations and enabling zero-shot transfer to unseen distortions. We evaluate three LALMs across three benchmarks and show that AnchorPrompt improves answer consistency in most tested conditions. Clean accuracy improves in six of nine model-benchmark pairs, with minimal impact on the remainder of 1.2% at most. Crucially, AnchorPrompt reduces hallucinations under severe audio corruption while keeping false refusals on clean audio rare. Finally, these consistency gains transfer to unseen perturbations, such as choice permutations and reverberation.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Mousavi, P., Ivry, A., Ravanelli, M., & Subakan, C. (2026). AnchorPrompt: Self-Distilled Soft Prompts for Robust Audio-Language Models. https://omanscience.com/en/articles/anchorprompt-self-distilled-soft-prompts-for-robust-audio-language-models

MLA 9

Mousavi, Pooneh, et al. "AnchorPrompt: Self-Distilled Soft Prompts for Robust Audio-Language Models." https://omanscience.com/en/articles/anchorprompt-self-distilled-soft-prompts-for-robust-audio-language-models.

Chicago (author–date)

Mousavi, Pooneh, Amir Ivry, Mirco Ravanelli, and Cem Subakan. 2026. "AnchorPrompt: Self-Distilled Soft Prompts for Robust Audio-Language Models." https://omanscience.com/en/articles/anchorprompt-self-distilled-soft-prompts-for-robust-audio-language-models.

Harvard

Mousavi, P., Ivry, A., Ravanelli, M. and Subakan, C. (2026) 'AnchorPrompt: Self-Distilled Soft Prompts for Robust Audio-Language Models', Available at: https://omanscience.com/en/articles/anchorprompt-self-distilled-soft-prompts-for-robust-audio-language-models.

Vancouver

Mousavi P, Ivry A, Ravanelli M, Subakan C. AnchorPrompt: Self-Distilled Soft Prompts for Robust Audio-Language Models. https://omanscience.com/en/articles/anchorprompt-self-distilled-soft-prompts-for-robust-audio-language-models

IEEE

P. Mousavi, A. Ivry, M. Ravanelli, and C. Subakan, "AnchorPrompt: Self-Distilled Soft Prompts for Robust Audio-Language Models," https://omanscience.com/en/articles/anchorprompt-self-distilled-soft-prompts-for-robust-audio-language-models.