Preprint Open access
CIPHER-MoE: Balancing Efficiency and Routing Fidelity in Trillion-Scale MoE Training
Mixture-of-Experts (MoE) has been widely adopted in recent large language model (LLM) architectures. However, scaling up MoE in LLM training introduces system-level challenges on training, where non-uniform token routing can lead to highly imbalanced workloads across experts and devices, further destabilizing the train …