Abstract
Deploying Mixture-of-Experts (MoE) models relies heavily on Expert Parallelism, which generates intense inter-GPU communication. Consequently, state-of-the-art inference systems require high-bandwidth, P2P interconnects (e.g., NVLink) in datacenter GPUs to handle massive token routing, making deployment prohibitively expensive. Consumer GPUs offer comparable compute power at significantly lower cost, promising to democratize MoE inference for individuals and enable privacy-preserving local deployments. However, their bandwidth-limited (only weak PCIe bus bandwidth) and host-mediated interconnects (no P2P support) introduce severe communication bottlenecks. We present CoMoE, a communication-efficient MoE inference system that resolves this mismatch through novel host-centric routing. Our key insight is that the unique communication topology provides the opportunity to elevate the host to an active routing hub, which can fundamentally reduce communication volume and eliminate global synchronization-induced stalls. Specifically, for token dispatch, we introduce host-backed token multicast to write shared tokens to the host exactly once, eliminating outbound transmission redundancy. For token combine, we propose a fine-grained, token-level aggregation mechanism using host staging buffers, which replaces rigid global synchronization and mitigates straggler effects. Evaluation on RTX 5090 GPUs shows that CoMoE improves inference throughput by up to 1.46x, approaching the performance of NVLink-capable A800 GPUs at only 23.4% of the hardware cost.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Fan, R., Zu, Y., Li, J., Hu, Q., Xinjun, Yang, Shu, J., & Lu, Y. (2026). Democratizing MoE inference on commodity GPUs with CoMoE. https://omanscience.com/en/articles/democratizing-moe-inference-on-commodity-gpus-with-comoe
MLA 9
Fan, Ruwen, et al. "Democratizing MoE inference on commodity GPUs with CoMoE." https://omanscience.com/en/articles/democratizing-moe-inference-on-commodity-gpus-with-comoe.
Chicago (author–date)
Fan, Ruwen, Yuezhi Zu, Junru Li, Qingda Hu, Xinjun, Yang, Jiwu Shu, and Youyou Lu. 2026. "Democratizing MoE inference on commodity GPUs with CoMoE." https://omanscience.com/en/articles/democratizing-moe-inference-on-commodity-gpus-with-comoe.
Harvard
Fan, R., Zu, Y., Li, J., Hu, Q., Xinjun, Yang, Shu, J. and Lu, Y. (2026) 'Democratizing MoE inference on commodity GPUs with CoMoE', Available at: https://omanscience.com/en/articles/democratizing-moe-inference-on-commodity-gpus-with-comoe.
Vancouver
Fan R, Zu Y, Li J, Hu Q, Xinjun, Yang, et al. Democratizing MoE inference on commodity GPUs with CoMoE. https://omanscience.com/en/articles/democratizing-moe-inference-on-commodity-gpus-with-comoe
IEEE
R. Fan, Y. Zu, J. Li, Q. Hu, Xinjun, Yang, J. Shu, and Y. Lu, "Democratizing MoE inference on commodity GPUs with CoMoE," https://omanscience.com/en/articles/democratizing-moe-inference-on-commodity-gpus-with-comoe.