Abstract

As Mixture-of-Experts (MoE) models scale toward hundreds of experts and higher top-$k$ routing, memory efficiency in distributed training becomes a critical bottleneck. Peak memory is dominated by the MoE block, not attention: every intermediate buffer in the MoE dispatch pipeline is individually scaled by top-k routing. The standard all-to-all dispatcher sends all routed tokens in a single collective step, requiring the full top-$k$-expanded buffer to be constructed at once. In this work, we propose RelayMoE, a ring-based MoE execution model that computes locally as expert weights or tokens circulate, avoiding full top-$k$-expanded dispatch buffers. RelayMoE selects between expert and token routing according to communication volume and overlaps transfers with computation. The ring structure naturally supports memory-efficient MoE recomputation during backward: each hop reconstructs expert intermediates, uses them to compute gradients, and releases them before the next hop. The saved memory supports longer sequences and larger batches, or retains more attention activations to reduce attention recomputation and improve training throughput. We evaluate RelayMoE on 30B$-$57B production MoE models and varied expert configurations. In single-layer MoE experiments, RelayMoE achieves a $2\times$ average speedup over Megatron-LM. In full-model training under the same GPU memory budget, it improves throughput by up to $2.02\times$ and extends the largest tested trainable sequence length by up to $2.85\times$.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Tarafder, A. K., Guasch-Martí, J., Kestor, G., & Ren, J. (2026). Memory-Efficient Expert Routing for Distributed MoE Training. https://omanscience.com/en/articles/memory-efficient-expert-routing-for-distributed-moe-training

MLA 9

Tarafder, Arnab Kanti, et al. "Memory-Efficient Expert Routing for Distributed MoE Training." https://omanscience.com/en/articles/memory-efficient-expert-routing-for-distributed-moe-training.

Chicago (author–date)

Tarafder, Arnab Kanti, Jaume Guasch-Martí, Gokcen Kestor, and Jie Ren. 2026. "Memory-Efficient Expert Routing for Distributed MoE Training." https://omanscience.com/en/articles/memory-efficient-expert-routing-for-distributed-moe-training.

Harvard

Tarafder, A. K., Guasch-Martí, J., Kestor, G. and Ren, J. (2026) 'Memory-Efficient Expert Routing for Distributed MoE Training', Available at: https://omanscience.com/en/articles/memory-efficient-expert-routing-for-distributed-moe-training.

Vancouver

Tarafder AK, Guasch-Martí J, Kestor G, Ren J. Memory-Efficient Expert Routing for Distributed MoE Training. https://omanscience.com/en/articles/memory-efficient-expert-routing-for-distributed-moe-training

IEEE

A. K. Tarafder, J. Guasch-Martí, G. Kestor, and J. Ren, "Memory-Efficient Expert Routing for Distributed MoE Training," https://omanscience.com/en/articles/memory-efficient-expert-routing-for-distributed-moe-training.