نسخة أولية وصول مفتوح
Attention Routing Stabilizes Early: Working-Set Inference for Recurrent Language Models
Recurrent-depth language models, such as looped Transformers, repeatedly apply shared network blocks to refine latent representations without generating explicit intermediate reasoning tokens. However, each step recomputes full attention over the entire context, repeating costly global routing. We study how attention r …