الباحثون

Anastasios Alexandridis

المنشورات 1

نسخة أولية وصول مفتوح

CommunityKV: Efficient Long-Context Decoding via Graph Partitioning

Scaling Transformers to long contexts is constrained by the quadratic cost of self-attention and the linear growth of key-value cache memory transfer. Sparse attention mitigates this by retrieving only relevant tokens, but current approaches either require large-scale training or, within the training-free regime, rely …

المؤلفون المشاركون