Abstract

Semi-structured pruning compresses large language models (LLMs) while keeping a regular sparse structure, but the prevailing N:M pattern fixes the same local sparsity ratio in every layer. Layer-adaptive sparsity allocation improves unstructured pruning, yet it has been reported to be less effective under N:M sparsity, leaving open whether adaptive allocation is of limited value for semi-structured pruning in general or only under the fine-grained N:M pattern. We examine this question with group-level sparsity, which partitions each weight matrix into regular groups, retains or prunes each group as a whole, and allows each layer's sparsity ratio to vary under a global budget. We propose GroupMask, which generates the group selectors of all layers with a lightweight hypernetwork, relaxes them with a Gumbel-Sigmoid parameterization and a straight-through estimator, and learns them through sparsity-budget regularization and self-distillation while keeping the pretrained weights frozen. On LLaMA-2-7B at 50% sparsity with the same $1\times256$ group size, learned layer-adaptive allocation reduces WikiText-2 perplexity from 10.02 to 8.30 and raises the average zero-shot accuracy from 0.455 to 0.496 relative to a uniform per-layer ratio. GroupMask obtains the lowest WikiText-2 perplexity on LLaMA-2-7B and the highest average zero-shot accuracy with Alpaca calibration among the evaluated baselines on five LLaMA and Qwen models. Our code is available at https://github.com/ZhengaoLi/GroupMask.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Li, Z., Li, S., Zhang, X., Jin, Y., Kestor, G., Zhang, Y., Zeng, Y., Ren, B., Zhang, C., & Gao, S. (2026). GroupMask: Layer-Adaptive Group-wise Sparsity for Semi-Structured LLM Pruning. https://omanscience.com/en/articles/groupmask-layer-adaptive-group-wise-sparsity-for-semi-structured-llm-pruning

MLA 9

Li, Zhengao, et al. "GroupMask: Layer-Adaptive Group-wise Sparsity for Semi-Structured LLM Pruning." https://omanscience.com/en/articles/groupmask-layer-adaptive-group-wise-sparsity-for-semi-structured-llm-pruning.

Chicago (author–date)

Li, Zhengao, Shuoqiu Li, Xiaofang Zhang, Yukai Jin, Gokcen Kestor, Yanfu Zhang, Yiming Zeng, Bin Ren, Chuxu Zhang, and Shangqian Gao. 2026. "GroupMask: Layer-Adaptive Group-wise Sparsity for Semi-Structured LLM Pruning." https://omanscience.com/en/articles/groupmask-layer-adaptive-group-wise-sparsity-for-semi-structured-llm-pruning.

Harvard

Li, Z., Li, S., Zhang, X., Jin, Y., Kestor, G., Zhang, Y., Zeng, Y., Ren, B., Zhang, C. and Gao, S. (2026) 'GroupMask: Layer-Adaptive Group-wise Sparsity for Semi-Structured LLM Pruning', Available at: https://omanscience.com/en/articles/groupmask-layer-adaptive-group-wise-sparsity-for-semi-structured-llm-pruning.

Vancouver

Li Z, Li S, Zhang X, Jin Y, Kestor G, Zhang Y, et al. GroupMask: Layer-Adaptive Group-wise Sparsity for Semi-Structured LLM Pruning. https://omanscience.com/en/articles/groupmask-layer-adaptive-group-wise-sparsity-for-semi-structured-llm-pruning

IEEE

Z. Li, S. Li, X. Zhang, Y. Jin, G. Kestor, Y. Zhang, Y. Zeng, B. Ren, C. Zhang, and S. Gao, "GroupMask: Layer-Adaptive Group-wise Sparsity for Semi-Structured LLM Pruning," https://omanscience.com/en/articles/groupmask-layer-adaptive-group-wise-sparsity-for-semi-structured-llm-pruning.