Abstract

Repeated queries over shared documents incur redundant encoding, while caching model states introduces persistent storage costs. Building on CoMem's intermediate-state interface, EncBank treats a pretrained LLM's lower layers as a reusable document encoder and compactly stores their outputs for an adapted upper-layer reader. A self-distilled suffix adapter is shared across storage precisions within each backbone, without quantization-specific retraining. Across five benchmark suites on three Qwen backbones spanning different sizes and full-attention and hybrid architectures, 4-bit storage keeps each reported benchmark aggregate within one score point of native-precision EncBank. In a fixed Qwen3-8B workload, it retains 28.1% of the native-precision persistent GPU store. Separate native-precision controls yield a 1.40x selected-pack prefill speedup over same-evidence, same-adapter text replay, at a 3.12-point RULER accuracy cost. A native-precision Qwen3.8-27B configuration also passes 70 of 89 Terminal-Bench 2.1 tasks. EncBank thus combines reusable computation with compact memory, while task fidelity and end-to-end benefits remain dependent on the workload, preparation costs, and reuse frequency.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Liu, H., Liu, C., Lin, C., Lamb, A., & Gao, M. (2026). Cache the Encoder Within:Compact, Reusable Memory across LLM Queries. https://omanscience.com/en/articles/cache-the-encoder-within-compact-reusable-memory-across-llm-queries

MLA 9

Liu, Hanzuo, et al. "Cache the Encoder Within:Compact, Reusable Memory across LLM Queries." https://omanscience.com/en/articles/cache-the-encoder-within-compact-reusable-memory-across-llm-queries.

Chicago (author–date)

Liu, Hanzuo, Chunyu Liu, Chaofan Lin, Alex Lamb, and Mingyu Gao. 2026. "Cache the Encoder Within:Compact, Reusable Memory across LLM Queries." https://omanscience.com/en/articles/cache-the-encoder-within-compact-reusable-memory-across-llm-queries.

Harvard

Liu, H., Liu, C., Lin, C., Lamb, A. and Gao, M. (2026) 'Cache the Encoder Within:Compact, Reusable Memory across LLM Queries', Available at: https://omanscience.com/en/articles/cache-the-encoder-within-compact-reusable-memory-across-llm-queries.

Vancouver

Liu H, Liu C, Lin C, Lamb A, Gao M. Cache the Encoder Within:Compact, Reusable Memory across LLM Queries. https://omanscience.com/en/articles/cache-the-encoder-within-compact-reusable-memory-across-llm-queries

IEEE

H. Liu, C. Liu, C. Lin, A. Lamb, and M. Gao, "Cache the Encoder Within:Compact, Reusable Memory across LLM Queries," https://omanscience.com/en/articles/cache-the-encoder-within-compact-reusable-memory-across-llm-queries.