Abstract
As Mixture-of-Experts (MoE) models scale toward hundreds of experts and higher top-$k$ routing, memory efficiency in distributed training becomes a critical bottleneck. Peak memory is dominated by the MoE block, not attention: every intermediate buffer in the MoE dispatch pipeline is individually scaled by top-k routing. The standard all-to-all dispatcher sends all routed tokens in a single collective step, requiring the full top-$k$-expanded buffer to be constructed at once. In this work, we propose RelayMoE, a ring-based MoE execution model that computes locally as expert weights or tokens circulate, avoiding full top-$k$-expanded dispatch buffers. RelayMoE selects between expert and token routing according to communication volume and overlaps transfers with computation. The ring structure naturally supports memory-efficient MoE recomputation during backward: each hop reconstructs expert intermediates, uses them to compute gradients, and releases them before the next hop. The saved memory supports longer sequences and larger batches, or retains more attention activations to reduce attention recomputation and improve training throughput. We evaluate RelayMoE on 30B$-$57B production MoE models and varied expert configurations. In single-layer MoE experiments, RelayMoE achieves a $2\times$ average speedup over Megatron-LM. In full-model training under the same GPU memory budget, it improves throughput by up to $2.02\times$ and extends the largest tested trainable sequence length by up to $2.85\times$.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Tarafder, A. K., Guasch-Martí, J., Kestor, G., & Ren, J. (2026). Memory-Efficient Expert Routing for Distributed MoE Training. https://omanscience.com/en/articles/memory-efficient-expert-routing-for-distributed-moe-training
MLA 9
Tarafder, Arnab Kanti, et al. "Memory-Efficient Expert Routing for Distributed MoE Training." https://omanscience.com/en/articles/memory-efficient-expert-routing-for-distributed-moe-training.
Chicago (author–date)
Tarafder, Arnab Kanti, Jaume Guasch-Martí, Gokcen Kestor, and Jie Ren. 2026. "Memory-Efficient Expert Routing for Distributed MoE Training." https://omanscience.com/en/articles/memory-efficient-expert-routing-for-distributed-moe-training.
Harvard
Tarafder, A. K., Guasch-Martí, J., Kestor, G. and Ren, J. (2026) 'Memory-Efficient Expert Routing for Distributed MoE Training', Available at: https://omanscience.com/en/articles/memory-efficient-expert-routing-for-distributed-moe-training.
Vancouver
Tarafder AK, Guasch-Martí J, Kestor G, Ren J. Memory-Efficient Expert Routing for Distributed MoE Training. https://omanscience.com/en/articles/memory-efficient-expert-routing-for-distributed-moe-training
IEEE
A. K. Tarafder, J. Guasch-Martí, G. Kestor, and J. Ren, "Memory-Efficient Expert Routing for Distributed MoE Training," https://omanscience.com/en/articles/memory-efficient-expert-routing-for-distributed-moe-training.