الباحثون

Jongwoo Ko

المنشورات 1

نسخة أولية وصول مفتوح

AttSVD:Prompt-Adaptive Low-Rank KV Cache Compression via Attention-Guided SVD

The key-value (KV) cache of autoregressive transformers grows linearly with context length and dominates memory at long context. Most training-free remedies evict low-importance tokens, an irreversible choice along the sequence axis. We instead keep every token and store it more cheaply along the "feature" axis. We the …

المؤلفون المشاركون