نسخة أولية وصول مفتوح
CacheRepair: Learning to Repair Cross-Chunk Context in RAG for KV Cache Fusion
Multi-document retrieval-augmented generation (RAG) requires a language model to process multiple retrieved text chunks before answering a question. Precomputing each chunk's KV cache independently and concatenating the caches when the chunks are retrieved can accelerate this step. However, the assembled cache lacks cr …