Abstract

Mixture-of-experts (MoE) layers increase model capacity without a proportional increase in per-example computation. However, conventional flat routers can yield imbalanced expert utilization and treat experts as an unstructured collection, whose indices carry no topological meaning. We introduce {\bf BRANCH-MoE}, a routing architecture that places \(E\) experts at the leaves of a binary decision tree of depth \(\log_2 E\). At each internal node the branching probability is centered on the arrival-weighted mean score of the traffic reaching that node. This mean is estimated using an exponential moving average, which promotes utilization of both child subtrees without an auxiliary load-balancing loss. We show that this moving-average estimate admits an explicit noise-lag trade-off. We prove that for linear node maps and log-concave arrival distributions, this mechanism prevents routing-mass collapse. We further establish that, under a frozen router, an expert's execution frequency controls its stochastic-gradient convergence rate, and that confident decisions near the root bound cross-device communication when experts are assigned to devices by tree prefix. We evaluate BRANCH-MoE against Switch softmax, DeepSeek-V3 dynamic-bias, Skywork logit-normalized, and deterministic hash routing on Criteo click-through-rate prediction, Forest Covertype, HIGGS, and YearPredictionMSD, using \(E=16\), top-\(4\) routing, and five random seeds. Our results show that hierarchical routing can preserve task quality and balanced utilization while inducing a topology that supports localized expert co-activation and reduced communication.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Fu, G., Javanmard, A., Bateni, M., & Mirrokni, V. (2026). BRANCH-MoE: Balance-Aware Tree Routing for Large Embedding Models. https://omanscience.com/en/articles/branch-moe-balance-aware-tree-routing-for-large-embedding-models

MLA 9

Fu, Gang, et al. "BRANCH-MoE: Balance-Aware Tree Routing for Large Embedding Models." https://omanscience.com/en/articles/branch-moe-balance-aware-tree-routing-for-large-embedding-models.

Chicago (author–date)

Fu, Gang, Adel Javanmard, MohammadHossein Bateni, and Vahab Mirrokni. 2026. "BRANCH-MoE: Balance-Aware Tree Routing for Large Embedding Models." https://omanscience.com/en/articles/branch-moe-balance-aware-tree-routing-for-large-embedding-models.

Harvard

Fu, G., Javanmard, A., Bateni, M. and Mirrokni, V. (2026) 'BRANCH-MoE: Balance-Aware Tree Routing for Large Embedding Models', Available at: https://omanscience.com/en/articles/branch-moe-balance-aware-tree-routing-for-large-embedding-models.

Vancouver

Fu G, Javanmard A, Bateni M, Mirrokni V. BRANCH-MoE: Balance-Aware Tree Routing for Large Embedding Models. https://omanscience.com/en/articles/branch-moe-balance-aware-tree-routing-for-large-embedding-models

IEEE

G. Fu, A. Javanmard, M. Bateni, and V. Mirrokni, "BRANCH-MoE: Balance-Aware Tree Routing for Large Embedding Models," https://omanscience.com/en/articles/branch-moe-balance-aware-tree-routing-for-large-embedding-models.