الباحثون

Pierre Ablin

المنشورات 3

نسخة أولية وصول مفتوح

KV-Lingo: Learning KV-Cache Translators with Distillation

Large language models represent context with a key-value (KV) cache. Caches are model-specific: for the same text, models with different architectures or weights produce incompatible representations. This makes it costly to switch models over a shared context: although the context has already been processed by one mode …

المؤلفون المشاركون