Abstract

The key-value (KV) cache of autoregressive transformers grows linearly with context length and dominates memory at long context. Most training-free remedies evict low-importance tokens, an irreversible choice along the sequence axis. We instead keep every token and store it more cheaply along the "feature" axis. We therefore propose AttSVD, a new "interpretable" low-rank compression whose basis is derived from each prompt's own attention geometry: an online, per-prompt truncated SVD that keeps only the directions attention actually reads, cutting persistent per-head KV memory in proportion to the retained rank. We propose two decode-time caching strategies, accumulating and streaming, for short and long generation regimes. Furthermore, we propose two refinements that make compression adaptive. A per-matrix energy rule sizes the logit space and the attention mass independently. An attention-aware basis truncates only in the spaces attention actually reads, preserving both the attention logits and the attention output. The same factors also provide free, per-head interpretability insights into the effective rank and the geometry attention consumes. Across multiple models, on both an agentic benchmark and the full LongBench suite AttSVD stays on par with the dense cache while using up to 50% of the KV-cache memory.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Abdali, S., Ko, J., & Cameron, P. (2026). AttSVD:Prompt-Adaptive Low-Rank KV Cache Compression via Attention-Guided SVD. https://omanscience.com/en/articles/attsvd-prompt-adaptive-low-rank-kv-cache-compression-via-attention-guided-svd

MLA 9

Abdali, Sara, et al. "AttSVD:Prompt-Adaptive Low-Rank KV Cache Compression via Attention-Guided SVD." https://omanscience.com/en/articles/attsvd-prompt-adaptive-low-rank-kv-cache-compression-via-attention-guided-svd.

Chicago (author–date)

Abdali, Sara, Jongwoo Ko, and Pashmina Cameron. 2026. "AttSVD:Prompt-Adaptive Low-Rank KV Cache Compression via Attention-Guided SVD." https://omanscience.com/en/articles/attsvd-prompt-adaptive-low-rank-kv-cache-compression-via-attention-guided-svd.

Harvard

Abdali, S., Ko, J. and Cameron, P. (2026) 'AttSVD:Prompt-Adaptive Low-Rank KV Cache Compression via Attention-Guided SVD', Available at: https://omanscience.com/en/articles/attsvd-prompt-adaptive-low-rank-kv-cache-compression-via-attention-guided-svd.

Vancouver

Abdali S, Ko J, Cameron P. AttSVD:Prompt-Adaptive Low-Rank KV Cache Compression via Attention-Guided SVD. https://omanscience.com/en/articles/attsvd-prompt-adaptive-low-rank-kv-cache-compression-via-attention-guided-svd

IEEE

S. Abdali, J. Ko, and P. Cameron, "AttSVD:Prompt-Adaptive Low-Rank KV Cache Compression via Attention-Guided SVD," https://omanscience.com/en/articles/attsvd-prompt-adaptive-low-rank-kv-cache-compression-via-attention-guided-svd.