الباحثون

Marco Cuturi

المنشورات 3

نسخة أولية وصول مفتوح

KV-Kaizen: Learning Context-Adaptive Cache Compression Choices

As the context size of text processed with an LLM grows, the size of KV caches can outstrip the memory allocated for the original model weights. This impacts LLM throughput negatively, since decoding is memory-bound and decode cost grows with cache size. Recent work alleviates this bottleneck by discarding the least re …

نسخة أولية وصول مفتوح

KV-Lingo: Learning KV-Cache Translators with Distillation

Large language models represent context with a key-value (KV) cache. Caches are model-specific: for the same text, models with different architectures or weights produce incompatible representations. This makes it costly to switch models over a shared context: although the context has already been processed by one mode …

المؤلفون المشاركون