Abstract
Large audio-language models (LALMs) are sensitive to input perturbations, such as noise, waveform corruption, and adversarial injections. We propose AnchorPrompt, an efficient adaptation method that keeps the model frozen and learns a single block of prompt vectors inserted at the decoder input, between the audio and question embeddings. We train these vectors through self-distillation over diverse audio and text perturbations. To improve answer consistency and mitigate hallucination, we use the model's prediction on the clean recording as the target for answerable inputs, and assign a refusal target when the audio lacks sufficient evidence to answer. Furthermore, AnchorPrompt is perturbation-agnostic at inference, requiring no prior detection of perturbations and enabling zero-shot transfer to unseen distortions. We evaluate three LALMs across three benchmarks and show that AnchorPrompt improves answer consistency in most tested conditions. Clean accuracy improves in six of nine model-benchmark pairs, with minimal impact on the remainder of 1.2% at most. Crucially, AnchorPrompt reduces hallucinations under severe audio corruption while keeping false refusals on clean audio rare. Finally, these consistency gains transfer to unseen perturbations, such as choice permutations and reverberation.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Mousavi, P., Ivry, A., Ravanelli, M., & Subakan, C. (2026). AnchorPrompt: Self-Distilled Soft Prompts for Robust Audio-Language Models. https://omanscience.com/en/articles/anchorprompt-self-distilled-soft-prompts-for-robust-audio-language-models
MLA 9
Mousavi, Pooneh, et al. "AnchorPrompt: Self-Distilled Soft Prompts for Robust Audio-Language Models." https://omanscience.com/en/articles/anchorprompt-self-distilled-soft-prompts-for-robust-audio-language-models.
Chicago (author–date)
Mousavi, Pooneh, Amir Ivry, Mirco Ravanelli, and Cem Subakan. 2026. "AnchorPrompt: Self-Distilled Soft Prompts for Robust Audio-Language Models." https://omanscience.com/en/articles/anchorprompt-self-distilled-soft-prompts-for-robust-audio-language-models.
Harvard
Mousavi, P., Ivry, A., Ravanelli, M. and Subakan, C. (2026) 'AnchorPrompt: Self-Distilled Soft Prompts for Robust Audio-Language Models', Available at: https://omanscience.com/en/articles/anchorprompt-self-distilled-soft-prompts-for-robust-audio-language-models.
Vancouver
Mousavi P, Ivry A, Ravanelli M, Subakan C. AnchorPrompt: Self-Distilled Soft Prompts for Robust Audio-Language Models. https://omanscience.com/en/articles/anchorprompt-self-distilled-soft-prompts-for-robust-audio-language-models
IEEE
P. Mousavi, A. Ivry, M. Ravanelli, and C. Subakan, "AnchorPrompt: Self-Distilled Soft Prompts for Robust Audio-Language Models," https://omanscience.com/en/articles/anchorprompt-self-distilled-soft-prompts-for-robust-audio-language-models.