Authors

Yuezhi Zu

Publications 1

Preprint Open access

Democratizing MoE inference on commodity GPUs with CoMoE

Ruwen Fan, Yuezhi Zu, Junru Li et al. · 2026

Deploying Mixture-of-Experts (MoE) models relies heavily on Expert Parallelism, which generates intense inter-GPU communication. Consequently, state-of-the-art inference systems require high-bandwidth, P2P interconnects (e.g., NVLink) in datacenter GPUs to handle massive token routing, making deployment prohibitively e …

Co-authors