Abstract

Large Language Models (LLMs) achieve strong performance across many domains, but their efficiency is limited by the quadratic cost of attention with respect to prompt length. Sparse attention reduces this cost by retaining only a small fraction of query-key interactions to approximate the full attention matrix. However, existing methods are trapped in a mathematically wrong view: they simply keep large scalar entries or high-mass regions of the attention matrix. This treats the attention matrix as a bag of values, ignoring that it is used as a structured matrix whose entries jointly determine the attention output through multiplication with value vectors. We argue that this is the core conceptual issue: sparse attention should be formulated as matrix approximation, not as blindly choosing the largest values from a bag of entries. Based on this view, we propose Matrix Approximation Sparse Attention (MASA). MASA replaces raw attention-mass ranking with a closed-form score that measures how much each sparse unit reduces matrix-product approximation error. As a theory-grounded plug-in correction, MASA can be added to existing sparse attention frameworks without changing their sparse kernels or budgets. Extensive experiments across multiple sparse attention methods, benchmarks, and LLM backbones show consistent accuracy gains, supporting both MASA and the matrix-approximation view of sparse attention.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Wan, F., Liu, X., Li, F., & Liu, Y. (2026). Sparse Attention Is Matrix Approximation, Not Choosing from a Bag of Values. https://omanscience.com/en/articles/sparse-attention-is-matrix-approximation-not-choosing-from-a-bag-of-values

MLA 9

Wan, Fang, et al. "Sparse Attention Is Matrix Approximation, Not Choosing from a Bag of Values." https://omanscience.com/en/articles/sparse-attention-is-matrix-approximation-not-choosing-from-a-bag-of-values.

Chicago (author–date)

Wan, Fang, Xufeng Liu, Fan Li, and Yi Liu. 2026. "Sparse Attention Is Matrix Approximation, Not Choosing from a Bag of Values." https://omanscience.com/en/articles/sparse-attention-is-matrix-approximation-not-choosing-from-a-bag-of-values.

Harvard

Wan, F., Liu, X., Li, F. and Liu, Y. (2026) 'Sparse Attention Is Matrix Approximation, Not Choosing from a Bag of Values', Available at: https://omanscience.com/en/articles/sparse-attention-is-matrix-approximation-not-choosing-from-a-bag-of-values.

Vancouver

Wan F, Liu X, Li F, Liu Y. Sparse Attention Is Matrix Approximation, Not Choosing from a Bag of Values. https://omanscience.com/en/articles/sparse-attention-is-matrix-approximation-not-choosing-from-a-bag-of-values

IEEE

F. Wan, X. Liu, F. Li, and Y. Liu, "Sparse Attention Is Matrix Approximation, Not Choosing from a Bag of Values," https://omanscience.com/en/articles/sparse-attention-is-matrix-approximation-not-choosing-from-a-bag-of-values.