نسخة أولية وصول مفتوح
GroupMask: Layer-Adaptive Group-wise Sparsity for Semi-Structured LLM Pruning
Semi-structured pruning compresses large language models (LLMs) while keeping a regular sparse structure, but the prevailing N:M pattern fixes the same local sparsity ratio in every layer. Layer-adaptive sparsity allocation improves unstructured pruning, yet it has been reported to be less effective under N:M sparsity, …