الملخص

تمت ترجمة أجزاء من هذه الصفحة آلياً وقد تحتوي على أخطاء.

A large audio language model (LALM) turns a minute of speech into 750-1,500 tokens and prefills every one. Image-token pruning often cuts after the language model's first layers, where image tokens draw little attention. Audio tokens draw much more attention there, and their ranking is still far from final, so audio needs a ranking before the language model runs. Surprisingly, the attention an audio token will receive across the language model is already linearly predictable from its encoder output, before the language model runs. A linear map, fitted in closed form without labels, predicts this all-layer attention ranking at $ρ\geq .69$ on eleven of thirteen LALMs. Our method, Triage, cuts audio tokens by this prediction and, on multiple choice, cuts again at layer 2, correcting the prediction with the attention observed there. Triage sets its compression without labels, under two budgets that limit how far its output may differ from the model's own full-audio output. At the conservative budget, its word error rate and accuracy stay within .04 of full audio. At the aggressive budget, Triage beats every baseline in all twelve transcription cases. On multiple choice, at 2.2-5x compression, it outperforms DART, the strongest baseline on average, by .043 in mean accuracy. Because it cuts before the language model, it raises the audio that fits in Qwen2.5-Omni-3B's context window from 21.8 to about 62 minutes. At its most compressive point, Triage lets one GPU serve 4x as many concurrent 5-minute streams of that model. Project page: https://audio-triage.github.io

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Park, K., Li, Y., & Qiu, L. (2026). انتباه رموز الصوت قابل للتنبؤ قبل أن يعمل نموذج اللغة. https://omanscience.com/ar/articles/audio-token-attention-is-predictable-before-the-language-model-runs

MLA 9

Park, Kyoungjun, et al. "انتباه رموز الصوت قابل للتنبؤ قبل أن يعمل نموذج اللغة." https://omanscience.com/ar/articles/audio-token-attention-is-predictable-before-the-language-model-runs.

شيكاغو (المؤلف–التاريخ)

Park, Kyoungjun, Yunzhe Li, and Lili Qiu. 2026. "انتباه رموز الصوت قابل للتنبؤ قبل أن يعمل نموذج اللغة." https://omanscience.com/ar/articles/audio-token-attention-is-predictable-before-the-language-model-runs.

هارفارد

Park, K., Li, Y. and Qiu, L. (2026) 'انتباه رموز الصوت قابل للتنبؤ قبل أن يعمل نموذج اللغة', Available at: https://omanscience.com/ar/articles/audio-token-attention-is-predictable-before-the-language-model-runs.

فانكوفر

Park K, Li Y, Qiu L. انتباه رموز الصوت قابل للتنبؤ قبل أن يعمل نموذج اللغة. https://omanscience.com/ar/articles/audio-token-attention-is-predictable-before-the-language-model-runs

IEEE

K. Park, Y. Li, and L. Qiu, "انتباه رموز الصوت قابل للتنبؤ قبل أن يعمل نموذج اللغة," https://omanscience.com/ar/articles/audio-token-attention-is-predictable-before-the-language-model-runs.