Abstract

Video world models achieve long-range temporal consistency by storing KV cache during generation, but the growing cache makes KV cache memory a major deployment bottleneck, which motivates low-bit quantization study for efficiency. Existing 2-bit KV cache quantization methods can achieve nearly lossless performance on VBench, however, when applied to video world models, we find they still cause severe temporal flickering and visual degradation. Meanwhile, deeper investigates show that Key quantization produces smaller reconstruction errors than Value, but surprisingly leads to larger output degradation. We trace this discrepancy to attention in video world models: Key perturbations can change the attention logits, and shift the temporal-spatial tokens selected by Queries. These observations motivate us to preserve attention logits and temporal-spatial token selection during KV cache quantization. To address this issue, we present QuantWM, a training-free 2-bit KV cache quantization framework for video world models. QuantWM introduces two complementary techniques to mitigate the attention shifts. Firstly, quantization-sensitivity-aware clustering (QSAC) jointly considers historical Query sensitivity and residual ranges to select INT2-friendly Key centroids, which reduces quantization errors in channels that are more critical to attention. In addition, principal-subspace attention compensation (PSAC) restores the remaining Key errors along the dominant Query subspace using low-rank projections, which provides a direct and efficient correction to stabilize attention logits. Experiments on LingBot-World-v2, HY-World 1.5, Matrix-Game-2, Longcat-Video and Causal-Forcing demonstrate that QuantWM significantly improves visual quality and temporal consistency, while outperforming existing methods across benchmarks with up to 6.20 KV cache memory compression and limited additional overhead.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Zhao, J., Hu, X., Yin, B., Jiang, J. P., Zhang, M., & Yan, S. (2026). QuantWM: Temporally Consistent 2-Bit KV Cache Quantization for Video World Models. https://omanscience.com/en/articles/quantwm-temporally-consistent-2-bit-kv-cache-quantization-for-video-world-models

MLA 9

Zhao, Jiaqi, et al. "QuantWM: Temporally Consistent 2-Bit KV Cache Quantization for Video World Models." https://omanscience.com/en/articles/quantwm-temporally-consistent-2-bit-kv-cache-quantization-for-video-world-models.

Chicago (author–date)

Zhao, Jiaqi, Xiaobin Hu, Bo Yin, Jun-Peng Jiang, Miao Zhang, and Shuicheng Yan. 2026. "QuantWM: Temporally Consistent 2-Bit KV Cache Quantization for Video World Models." https://omanscience.com/en/articles/quantwm-temporally-consistent-2-bit-kv-cache-quantization-for-video-world-models.

Harvard

Zhao, J., Hu, X., Yin, B., Jiang, J. P., Zhang, M. and Yan, S. (2026) 'QuantWM: Temporally Consistent 2-Bit KV Cache Quantization for Video World Models', Available at: https://omanscience.com/en/articles/quantwm-temporally-consistent-2-bit-kv-cache-quantization-for-video-world-models.

Vancouver

Zhao J, Hu X, Yin B, Jiang JP, Zhang M, Yan S. QuantWM: Temporally Consistent 2-Bit KV Cache Quantization for Video World Models. https://omanscience.com/en/articles/quantwm-temporally-consistent-2-bit-kv-cache-quantization-for-video-world-models

IEEE

J. Zhao, X. Hu, B. Yin, J. P. Jiang, M. Zhang, and S. Yan, "QuantWM: Temporally Consistent 2-Bit KV Cache Quantization for Video World Models," https://omanscience.com/en/articles/quantwm-temporally-consistent-2-bit-kv-cache-quantization-for-video-world-models.