نسخة أولية وصول مفتوح
Periodic Weak Spots: Phase Sensitivity from Chunked KV-Cache Compression
Chunked KV-cache compression reduces the memory and attention costs of long-context inference by compressing windows of consecutive tokens into fewer cache entries at a fixed stride. Such compression also introduces a new positional coordinate: a token's phase, or its position relative to compression-window boundaries. …