Abstract

Deploying Mixture-of-Experts (MoE) models relies heavily on Expert Parallelism, which generates intense inter-GPU communication. Consequently, state-of-the-art inference systems require high-bandwidth, P2P interconnects (e.g., NVLink) in datacenter GPUs to handle massive token routing, making deployment prohibitively expensive. Consumer GPUs offer comparable compute power at significantly lower cost, promising to democratize MoE inference for individuals and enable privacy-preserving local deployments. However, their bandwidth-limited (only weak PCIe bus bandwidth) and host-mediated interconnects (no P2P support) introduce severe communication bottlenecks. We present CoMoE, a communication-efficient MoE inference system that resolves this mismatch through novel host-centric routing. Our key insight is that the unique communication topology provides the opportunity to elevate the host to an active routing hub, which can fundamentally reduce communication volume and eliminate global synchronization-induced stalls. Specifically, for token dispatch, we introduce host-backed token multicast to write shared tokens to the host exactly once, eliminating outbound transmission redundancy. For token combine, we propose a fine-grained, token-level aggregation mechanism using host staging buffers, which replaces rigid global synchronization and mitigates straggler effects. Evaluation on RTX 5090 GPUs shows that CoMoE improves inference throughput by up to 1.46x, approaching the performance of NVLink-capable A800 GPUs at only 23.4% of the hardware cost.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Fan, R., Zu, Y., Li, J., Hu, Q., Xinjun, Yang, Shu, J., & Lu, Y. (2026). Democratizing MoE inference on commodity GPUs with CoMoE. https://omanscience.com/en/articles/democratizing-moe-inference-on-commodity-gpus-with-comoe

MLA 9

Fan, Ruwen, et al. "Democratizing MoE inference on commodity GPUs with CoMoE." https://omanscience.com/en/articles/democratizing-moe-inference-on-commodity-gpus-with-comoe.

Chicago (author–date)

Fan, Ruwen, Yuezhi Zu, Junru Li, Qingda Hu, Xinjun, Yang, Jiwu Shu, and Youyou Lu. 2026. "Democratizing MoE inference on commodity GPUs with CoMoE." https://omanscience.com/en/articles/democratizing-moe-inference-on-commodity-gpus-with-comoe.

Harvard

Fan, R., Zu, Y., Li, J., Hu, Q., Xinjun, Yang, Shu, J. and Lu, Y. (2026) 'Democratizing MoE inference on commodity GPUs with CoMoE', Available at: https://omanscience.com/en/articles/democratizing-moe-inference-on-commodity-gpus-with-comoe.

Vancouver

Fan R, Zu Y, Li J, Hu Q, Xinjun, Yang, et al. Democratizing MoE inference on commodity GPUs with CoMoE. https://omanscience.com/en/articles/democratizing-moe-inference-on-commodity-gpus-with-comoe

IEEE

R. Fan, Y. Zu, J. Li, Q. Hu, Xinjun, Yang, J. Shu, and Y. Lu, "Democratizing MoE inference on commodity GPUs with CoMoE," https://omanscience.com/en/articles/democratizing-moe-inference-on-commodity-gpus-with-comoe.