نسخة أولية وصول مفتوح
Audio Token Attention Is Predictable Before the Language Model Runs
A large audio language model (LALM) turns a minute of speech into 750-1,500 tokens and prefills every one. Image-token pruning often cuts after the language model's first layers, where image tokens draw little attention. Audio tokens draw much more attention there, and their ranking is still far from final, so audio ne …