الباحثون

Qingyan Meng

المنشورات 1

نسخة أولية وصول مفتوح

Looping Beyond Twice: A Scalable Recipe for Looped Mixture-of-Experts

Looped Transformers introduce recurrent depth as a new scaling axis for LLMs: by repeatedly applying shared Transformer blocks, they increase effective depth without increasing parameter count. However, the benefits of looping remain unclear for large MoE LLMs under FLOPs-matched comparisons. The main reason is that th …

المؤلفون المشاركون