Authors

Sungwon Kim

Publications 2

Preprint Open access

VFold: Symmetry-Aware Cross-Layer Value Cache Compression

While caching key-value (KV) states accelerates Large Language Model (LLM) decoding, this cache can dominate memory usage at long context lengths. One solution is to compress this memory by exploiting inter-layer cache similarities. However, most existing techniques necessitate architectural changes to LLMs and incur s …

Co-authors