الملخص

As Mixture-of-Experts (MoE) models scale toward hundreds of experts and higher top-$k$ routing, memory efficiency in distributed training becomes a critical bottleneck. Peak memory is dominated by the MoE block, not attention: every intermediate buffer in the MoE dispatch pipeline is individually scaled by top-k routing. The standard all-to-all dispatcher sends all routed tokens in a single collective step, requiring the full top-$k$-expanded buffer to be constructed at once. In this work, we propose RelayMoE, a ring-based MoE execution model that computes locally as expert weights or tokens circulate, avoiding full top-$k$-expanded dispatch buffers. RelayMoE selects between expert and token routing according to communication volume and overlaps transfers with computation. The ring structure naturally supports memory-efficient MoE recomputation during backward: each hop reconstructs expert intermediates, uses them to compute gradients, and releases them before the next hop. The saved memory supports longer sequences and larger batches, or retains more attention activations to reduce attention recomputation and improve training throughput. We evaluate RelayMoE on 30B$-$57B production MoE models and varied expert configurations. In single-layer MoE experiments, RelayMoE achieves a $2\times$ average speedup over Megatron-LM. In full-model training under the same GPU memory budget, it improves throughput by up to $2.02\times$ and extends the largest tested trainable sequence length by up to $2.85\times$.

الكلمات المفتاحية

الموضوع

بيانات النشر

المجلة
غير متاح
وصول مفتوح
وصول مفتوح أخضر

اقتبس هذه المقالة

APA 7

Tarafder, A. K., Guasch-Martí, J., Kestor, G., & Ren, J. (2026). Memory-Efficient Expert Routing for Distributed MoE Training. https://omanscience.com/ar/articles/memory-efficient-expert-routing-for-distributed-moe-training

MLA 9

Tarafder, Arnab Kanti, et al. "Memory-Efficient Expert Routing for Distributed MoE Training." https://omanscience.com/ar/articles/memory-efficient-expert-routing-for-distributed-moe-training.

شيكاغو (المؤلف–التاريخ)

Tarafder, Arnab Kanti, Jaume Guasch-Martí, Gokcen Kestor, and Jie Ren. 2026. "Memory-Efficient Expert Routing for Distributed MoE Training." https://omanscience.com/ar/articles/memory-efficient-expert-routing-for-distributed-moe-training.

هارفارد

Tarafder, A. K., Guasch-Martí, J., Kestor, G. and Ren, J. (2026) 'Memory-Efficient Expert Routing for Distributed MoE Training', Available at: https://omanscience.com/ar/articles/memory-efficient-expert-routing-for-distributed-moe-training.

فانكوفر

Tarafder AK, Guasch-Martí J, Kestor G, Ren J. Memory-Efficient Expert Routing for Distributed MoE Training. https://omanscience.com/ar/articles/memory-efficient-expert-routing-for-distributed-moe-training

IEEE

A. K. Tarafder, J. Guasch-Martí, G. Kestor, and J. Ren, "Memory-Efficient Expert Routing for Distributed MoE Training," https://omanscience.com/ar/articles/memory-efficient-expert-routing-for-distributed-moe-training.