الملخص

KV-cache quantization and linear attention are two representative approaches to tackling the storage and computational costs of Transformers. KV-cache quantization compresses individual KV entries into discrete codes but retains all entries, whereas linear attention recurrently aggregates multiple historical KV contributions into a fixed-size continuous state but can introduce interference. This contrast raises the question of whether per-KV compression and multi-KV aggregation can be bridged within a single mechanism for efficient attention. We identify RAM-Net as such a bridge through soft assignments over a discrete address space. These assignments determine recurrent updates to the continuous slot state associated with each address. Under a restricted RAM-Net construction, we prove that soft address assignments extend hard quantized matching to a separable read-write overlap that locally approximates full-attention similarity and supports recurrent aggregation. These connections further enable Transformer-to-RAM-Net weight migration through a new path based on a soft-quantized intermediate construction. Across nine pretrained Transformer models from 0.3B to 7B parameters, RAM-Net recovers an average of 87.1% of the teachers' accuracy gains over random guessing across six commonsense and knowledge tasks using only a 500M-token budget per model.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Xiao, K., Dong, L., Li, H., & Xing, G. (2026). Bridging KV-Cache Quantization and Linear Attention: From Theory to Pretrained Weight Migration. https://omanscience.com/ar/articles/bridging-kv-cache-quantization-and-linear-attention-from-theory-to-pretrained-weight-migration

MLA 9

Xiao, Kaicheng, et al. "Bridging KV-Cache Quantization and Linear Attention: From Theory to Pretrained Weight Migration." https://omanscience.com/ar/articles/bridging-kv-cache-quantization-and-linear-attention-from-theory-to-pretrained-weight-migration.

شيكاغو (المؤلف–التاريخ)

Xiao, Kaicheng, Liran Dong, Haotian Li, and Guoliang Xing. 2026. "Bridging KV-Cache Quantization and Linear Attention: From Theory to Pretrained Weight Migration." https://omanscience.com/ar/articles/bridging-kv-cache-quantization-and-linear-attention-from-theory-to-pretrained-weight-migration.

هارفارد

Xiao, K., Dong, L., Li, H. and Xing, G. (2026) 'Bridging KV-Cache Quantization and Linear Attention: From Theory to Pretrained Weight Migration', Available at: https://omanscience.com/ar/articles/bridging-kv-cache-quantization-and-linear-attention-from-theory-to-pretrained-weight-migration.

فانكوفر

Xiao K, Dong L, Li H, Xing G. Bridging KV-Cache Quantization and Linear Attention: From Theory to Pretrained Weight Migration. https://omanscience.com/ar/articles/bridging-kv-cache-quantization-and-linear-attention-from-theory-to-pretrained-weight-migration

IEEE

K. Xiao, L. Dong, H. Li, and G. Xing, "Bridging KV-Cache Quantization and Linear Attention: From Theory to Pretrained Weight Migration," https://omanscience.com/ar/articles/bridging-kv-cache-quantization-and-linear-attention-from-theory-to-pretrained-weight-migration.