Authors

Siyuan Zhang

Publications 3

Preprint Open access

Hybrid Latent Attention for Looped Language Models

Yuhan Chen, Siyuan Zhang, Nan Wang et al. · 2026

Looped language models apply the same stack of layers T times to each token, which deepens the model without adding parameters but multiplies its key-value (KV) cache by T. The larger cache limits how many sequences a GPU can decode at once and slows each decoding step, which reads the whole cache. We propose Hybrid La …

Co-authors