الباحثون

Gokcen Kestor

المنشورات 2

نسخة أولية وصول مفتوح

Memory-Efficient Expert Routing for Distributed MoE Training

As Mixture-of-Experts (MoE) models scale toward hundreds of experts and higher top-$k$ routing, memory efficiency in distributed training becomes a critical bottleneck. Peak memory is dominated by the MoE block, not attention: every intermediate buffer in the MoE dispatch pipeline is individually scaled by top-k routin …

نسخة أولية وصول مفتوح

GroupMask: Layer-Adaptive Group-wise Sparsity for Semi-Structured LLM Pruning

Zhengao Li, Shuoqiu Li, Xiaofang Zhang وآخرون · 2026

Semi-structured pruning compresses large language models (LLMs) while keeping a regular sparse structure, but the prevailing N:M pattern fixes the same local sparsity ratio in every layer. Layer-adaptive sparsity allocation improves unstructured pruning, yet it has been reported to be less effective under N:M sparsity, …

المؤلفون المشاركون