الباحثون

Xiaoqing Tong

المنشورات 1

نسخة أولية وصول مفتوح

How to Loop MoE: Flatten the Experts, Untie the Attention

Shouren Wang, Chuang Ma, Mohsen Hariri وآخرون · 2026

Looped Transformers reuse one block of layers several times: by spending extra computation they push a model of fixed size further, and so use its parameters more fully; while sparse mixture-of-experts (MoE) models activate only a few of many experts for each token. Looped MoE bridges these two design philosophies and …

المؤلفون المشاركون