Abstract
Activation sparsity speeds up large language model (LLM) inference by setting unimportant activations to zero so that the corresponding computations can be skipped. Existing training-free methods, however, make different trade-offs: threshold-based methods such as TEAL adapt the sparsity level to each token but do not tightly control the realised sparsity, while TopK-based methods such as WINA enforce a fixed sparsity level but use the same sparsity budget for every token. Both also apply the same budget across transformer blocks, despite large differences in block sensitivity. We introduce TopK-Guided, a training-free method that addresses both limitations by combining bounded token-level sparsity adaptation with sensitivity-aware block-level budget allocation. Across Llama-2 and Llama-3 models, TopK-Guided consistently improves perplexity and downstream accuracy over TEAL and WINA while preserving essentially the same sparsitydependent projection compute as WINA, with the largest gains at high sparsity. Ablations show that both components provide complementary improvements.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Agarwalla, M., & Lin, C. J. (2026). TopK-Guided: Adaptive, Budget-Aware Activation Sparsity for Efficient LLM Inference. https://omanscience.com/en/articles/topk-guided-adaptive-budget-aware-activation-sparsity-for-efficient-llm-inference
MLA 9
Agarwalla, Mukund, and Chih-Jen Lin. "TopK-Guided: Adaptive, Budget-Aware Activation Sparsity for Efficient LLM Inference." https://omanscience.com/en/articles/topk-guided-adaptive-budget-aware-activation-sparsity-for-efficient-llm-inference.
Chicago (author–date)
Agarwalla, Mukund, and Chih-Jen Lin. 2026. "TopK-Guided: Adaptive, Budget-Aware Activation Sparsity for Efficient LLM Inference." https://omanscience.com/en/articles/topk-guided-adaptive-budget-aware-activation-sparsity-for-efficient-llm-inference.
Harvard
Agarwalla, M. and Lin, C. J. (2026) 'TopK-Guided: Adaptive, Budget-Aware Activation Sparsity for Efficient LLM Inference', Available at: https://omanscience.com/en/articles/topk-guided-adaptive-budget-aware-activation-sparsity-for-efficient-llm-inference.
Vancouver
Agarwalla M, Lin CJ. TopK-Guided: Adaptive, Budget-Aware Activation Sparsity for Efficient LLM Inference. https://omanscience.com/en/articles/topk-guided-adaptive-budget-aware-activation-sparsity-for-efficient-llm-inference
IEEE
M. Agarwalla, and C. J. Lin, "TopK-Guided: Adaptive, Budget-Aware Activation Sparsity for Efficient LLM Inference," https://omanscience.com/en/articles/topk-guided-adaptive-budget-aware-activation-sparsity-for-efficient-llm-inference.