Abstract

A common small-model deployment runs one shared backbone with several LoRA specialists that answer over the same context. Serving them naively re-prefills that shared context once per specialist. We study a narrow, practical question: for already-trained standard LoRA adapters -- not adapters retrained for cache compatibility -- how much task quality is preserved if the backbone's prefill KV cache is computed once and reused across specialists, and what does that buy in serving cost? On a Qwen3-1.7B backbone with two adapters (extractive QA on HotpotQA, arithmetic reasoning on GSM8K), we sweep the boundary at which the specialist takes over from the reused base cache and measure paired quality differences and serving cost. Full-prefix reuse had the lowest prefill cost and a small quality difference on held-out GSM8K (Delta = -4.6 EM at a 160-token budget; -3.0 at 320 tokens; -0.8 under a second training seed -- all favoring native, only the first excluding zero, and the magnitude not consistent). Partial recomputation provided no demonstrated advantage. Neither quality equivalence nor a general boundary-selection rule is established. We also report a closed-form ridge KV translator that did not beat direct reuse, and specialist-dependence contrasts whose intervals all include zero. The measured serving benefit is warm-cache time-to-first-token, which grows with context (~16x at 8K); two-branch peak memory was only 12% lower and, on inspection, the prefix was never physically shared across branches -- this implementation reuses KV values but copies their storage, so shared-cache memory savings are not achieved.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Rajput, D. (2026). Shared-Prefix KV Reuse Across Standard LoRA Adapters: Quality and Serving Tradeoffs. https://omanscience.com/en/articles/shared-prefix-kv-reuse-across-standard-lora-adapters-quality-and-serving-tradeoffs

MLA 9

Rajput, Dushyant. "Shared-Prefix KV Reuse Across Standard LoRA Adapters: Quality and Serving Tradeoffs." https://omanscience.com/en/articles/shared-prefix-kv-reuse-across-standard-lora-adapters-quality-and-serving-tradeoffs.

Chicago (author–date)

Rajput, Dushyant. 2026. "Shared-Prefix KV Reuse Across Standard LoRA Adapters: Quality and Serving Tradeoffs." https://omanscience.com/en/articles/shared-prefix-kv-reuse-across-standard-lora-adapters-quality-and-serving-tradeoffs.

Harvard

Rajput, D. (2026) 'Shared-Prefix KV Reuse Across Standard LoRA Adapters: Quality and Serving Tradeoffs', Available at: https://omanscience.com/en/articles/shared-prefix-kv-reuse-across-standard-lora-adapters-quality-and-serving-tradeoffs.

Vancouver

Rajput D. Shared-Prefix KV Reuse Across Standard LoRA Adapters: Quality and Serving Tradeoffs. https://omanscience.com/en/articles/shared-prefix-kv-reuse-across-standard-lora-adapters-quality-and-serving-tradeoffs

IEEE

D. Rajput, "Shared-Prefix KV Reuse Across Standard LoRA Adapters: Quality and Serving Tradeoffs," https://omanscience.com/en/articles/shared-prefix-kv-reuse-across-standard-lora-adapters-quality-and-serving-tradeoffs.