الملخص
Scaling Transformers to long contexts is constrained by the quadratic cost of self-attention and the linear growth of key-value cache memory transfer. Sparse attention mitigates this by retrieving only relevant tokens, but current approaches either require large-scale training or, within the training-free regime, rely on semantically coarse heuristics or expensive clustering that is difficult to update efficiently during decoding. We introduce CommunityKV, a framework that formulates sparse attention as a community detection problem. CommunityKV constructs a token graph from the $QK^T$ scores already computed during standard prefill, and partitions the graph into communities to enable retrieval of semantically coherent token groups. A local update rule assigns newly generated tokens to communities in constant time, enabling sparse retrieval throughout streaming decoding without global re-partitioning. We evaluate CommunityKV on Qwen3 and Llama-3.1 models across three long-context benchmarks. With one graph per query head, CommunityKV delivers up to $1.25\times$ the end-to-end generation throughput of dense attention, while query-group graph aggregation yields up to $1.71\times$ with comparable accuracy.
الكلمات المفتاحية
الموضوع
بيانات النشر
- المجلة
- غير متاح
- وصول مفتوح
- وصول مفتوح أخضر
اقتبس هذه المقالة
APA 7
McKenna, J., Alexandridis, A., Susanj, N., & Liu, J. (2026). CommunityKV: Efficient Long-Context Decoding via Graph Partitioning. https://omanscience.com/ar/articles/communitykv-efficient-long-context-decoding-via-graph-partitioning
MLA 9
McKenna, Joe, et al. "CommunityKV: Efficient Long-Context Decoding via Graph Partitioning." https://omanscience.com/ar/articles/communitykv-efficient-long-context-decoding-via-graph-partitioning.
شيكاغو (المؤلف–التاريخ)
McKenna, Joe, Anastasios Alexandridis, Nathan Susanj, and Jing Liu. 2026. "CommunityKV: Efficient Long-Context Decoding via Graph Partitioning." https://omanscience.com/ar/articles/communitykv-efficient-long-context-decoding-via-graph-partitioning.
هارفارد
McKenna, J., Alexandridis, A., Susanj, N. and Liu, J. (2026) 'CommunityKV: Efficient Long-Context Decoding via Graph Partitioning', Available at: https://omanscience.com/ar/articles/communitykv-efficient-long-context-decoding-via-graph-partitioning.
فانكوفر
McKenna J, Alexandridis A, Susanj N, Liu J. CommunityKV: Efficient Long-Context Decoding via Graph Partitioning. https://omanscience.com/ar/articles/communitykv-efficient-long-context-decoding-via-graph-partitioning
IEEE
J. McKenna, A. Alexandridis, N. Susanj, and J. Liu, "CommunityKV: Efficient Long-Context Decoding via Graph Partitioning," https://omanscience.com/ar/articles/communitykv-efficient-long-context-decoding-via-graph-partitioning.