Authors

Xingyu Zhu

Publications 6

Preprint Open access

BARQ: Balanced Codebook Refinement for Low-Bit LLM Quantization

Chenhang Cui, Xu Xie, Linrui Xu et al. · 2026

As large language models (LLMs) grow in parameter count, model storage and parameter memory traffic have become major bottlenecks to efficient deployment. Codebook-based weight quantization reduces these costs, but imbalanced nearest-codeword assignments during fitting can leave some codewords insufficiently updated, l …

Preprint Open access

Periodic Weak Spots: Phase Sensitivity from Chunked KV-Cache Compression

Xingyu Zhu, Pu, Yi et al. · 2026

Chunked KV-cache compression reduces the memory and attention costs of long-context inference by compressing windows of consecutive tokens into fewer cache entries at a fixed stride. Such compression also introduces a new positional coordinate: a token's phase, or its position relative to compression-window boundaries. …

Co-authors