الباحثون

Jaume Guasch-Martí

المنشورات 1

نسخة أولية وصول مفتوح

Memory-Efficient Expert Routing for Distributed MoE Training

As Mixture-of-Experts (MoE) models scale toward hundreds of experts and higher top-$k$ routing, memory efficiency in distributed training becomes a critical bottleneck. Peak memory is dominated by the MoE block, not attention: every intermediate buffer in the MoE dispatch pipeline is individually scaled by top-k routin …

المؤلفون المشاركون