الباحثون

Changcheng Li

المنشورات 1

نسخة أولية وصول مفتوح

HERO-MoE: Historical Expert Routing with Scale-Preserving Fusion

Junxiang Qiu, Zhengsu Chen, Xinting Hu وآخرون · 2026

Mixture-of-Experts (MoE) architectures have become a standard way to scale model capacity while keeping computation sparse, yet routing remains a key determinant of MoE quality and training behavior. Prior empirical studies suggest that MoE routing reflects input semantics and upstream computation across depth, but sta …

المؤلفون المشاركون