نسخة أولية وصول مفتوح
ResidualQuant: KV Cache Quantization for Looped Transformers with 2-Bit Residuals
Looped Transformers improve parameter efficiency by repeatedly applying shared Transformer blocks over multiple recurrent loops, increasing computational depth without increasing the parameter count. However, KV cache memory still scales with the number of loops, becoming a key memory bottleneck that limits batch size …