الملخص
Large audio-language models (LALMs) exploit multimodal evidence, yet task-irrelevant audio can alter text-reasoning decisions when listening is unnecessary. Aggregate Accuracy can hide this paired drift because audio-induced repairs and damages may cancel. Paired drift analysis and targeted interventions identify architecture-specific, intervention-sensitive late audio pathways as actionable control points. We introduce ICAP-Gate, which applies mechanism-guided, task-conditioned control to each model's pathway. Across four LALMs, two reasoning benchmarks, and environmental-sound and natural-speech interference, ICAP-Gate has lower point estimates for Influence Rate and Answer Flip than ungated inference in all 16 full-split model--condition evaluations. Fixed suppression degrades automatic speech recognition (ASR) across all four models, whereas ICAP-Gate matches ungated ASR performance by preserving the pathway for explicit audio-demand instructions. ICAP-Gate has lower paired-drift point estimates than mitigation prompting in all four evaluated settings and provides competitive stabilization relative to eight-sample Self-Consistency while using one generation per query; in controlled ARC measurements, Self-Consistency incurs $7.0$--$9.2\times$ ungated latency. These results establish selective modality influence control as a design principle for robust multimodal reasoning.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
Sun, Y., Xu, K., & Dou, Y. (2026). Selective Listening: Mechanism-Guided Control of Audio Influence in Large Audio-Language Models. https://omanscience.com/ar/articles/selective-listening-mechanism-guided-control-of-audio-influence-in-large-audio-language-models
MLA 9
Sun, Yulin, et al. "Selective Listening: Mechanism-Guided Control of Audio Influence in Large Audio-Language Models." https://omanscience.com/ar/articles/selective-listening-mechanism-guided-control-of-audio-influence-in-large-audio-language-models.
شيكاغو (المؤلف–التاريخ)
Sun, Yulin, Kele Xu, and Yong Dou. 2026. "Selective Listening: Mechanism-Guided Control of Audio Influence in Large Audio-Language Models." https://omanscience.com/ar/articles/selective-listening-mechanism-guided-control-of-audio-influence-in-large-audio-language-models.
هارفارد
Sun, Y., Xu, K. and Dou, Y. (2026) 'Selective Listening: Mechanism-Guided Control of Audio Influence in Large Audio-Language Models', Available at: https://omanscience.com/ar/articles/selective-listening-mechanism-guided-control-of-audio-influence-in-large-audio-language-models.
فانكوفر
Sun Y, Xu K, Dou Y. Selective Listening: Mechanism-Guided Control of Audio Influence in Large Audio-Language Models. https://omanscience.com/ar/articles/selective-listening-mechanism-guided-control-of-audio-influence-in-large-audio-language-models
IEEE
Y. Sun, K. Xu, and Y. Dou, "Selective Listening: Mechanism-Guided Control of Audio Influence in Large Audio-Language Models," https://omanscience.com/ar/articles/selective-listening-mechanism-guided-control-of-audio-influence-in-large-audio-language-models.