نسخة أولية وصول مفتوح
MoRA: MoE Pruning via Router Bias Learning and Expert Approximation
Mixture-of-Experts (MoE) models enable parameter scaling with limited per-token computation by activating only a small subset of experts for each token, but deploying them still requires loading the complete expert pool into memory. Structured expert pruning can effectively reduce the memory usage by removing experts. …