نسخة أولية وصول مفتوح
Smoothing the Top-k Exposure Boundary for Sparse Mixture-of-Experts
Sparse Mixture-of-Experts models scale parameter capacity efficiently while maintaining a fixed compute budget per token. However, traditional training paradigms enforce a static choice of top-$k$ experts, which converts a continuous routing distribution into a rigid step function. This constraint introduces a brittle …