نسخة أولية وصول مفتوح
Sparse Attention Is Matrix Approximation, Not Choosing from a Bag of Values
Large Language Models (LLMs) achieve strong performance across many domains, but their efficiency is limited by the quadratic cost of attention with respect to prompt length. Sparse attention reduces this cost by retaining only a small fraction of query-key interactions to approximate the full attention matrix. However …