نسخة أولية وصول مفتوح
DivMoE: Fine-Grained MoE Upcycling via Cross-Domain Expert Composition
Mixture-of-Experts (MoE) architectures have become essential for scaling large language models, with recent work demonstrating the benefits of fine-grained expert designs. Training such models from scratch is expensive, and sparse upcycling from pre-trained dense models is an attractive alternative. However, we identif …