Abstract

Activation sparsity speeds up large language model (LLM) inference by setting unimportant activations to zero so that the corresponding computations can be skipped. Existing training-free methods, however, make different trade-offs: threshold-based methods such as TEAL adapt the sparsity level to each token but do not tightly control the realised sparsity, while TopK-based methods such as WINA enforce a fixed sparsity level but use the same sparsity budget for every token. Both also apply the same budget across transformer blocks, despite large differences in block sensitivity. We introduce TopK-Guided, a training-free method that addresses both limitations by combining bounded token-level sparsity adaptation with sensitivity-aware block-level budget allocation. Across Llama-2 and Llama-3 models, TopK-Guided consistently improves perplexity and downstream accuracy over TEAL and WINA while preserving essentially the same sparsitydependent projection compute as WINA, with the largest gains at high sparsity. Ablations show that both components provide complementary improvements.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Agarwalla, M., & Lin, C. J. (2026). TopK-Guided: Adaptive, Budget-Aware Activation Sparsity for Efficient LLM Inference. https://omanscience.com/en/articles/topk-guided-adaptive-budget-aware-activation-sparsity-for-efficient-llm-inference

MLA 9

Agarwalla, Mukund, and Chih-Jen Lin. "TopK-Guided: Adaptive, Budget-Aware Activation Sparsity for Efficient LLM Inference." https://omanscience.com/en/articles/topk-guided-adaptive-budget-aware-activation-sparsity-for-efficient-llm-inference.

Chicago (author–date)

Agarwalla, Mukund, and Chih-Jen Lin. 2026. "TopK-Guided: Adaptive, Budget-Aware Activation Sparsity for Efficient LLM Inference." https://omanscience.com/en/articles/topk-guided-adaptive-budget-aware-activation-sparsity-for-efficient-llm-inference.

Harvard

Agarwalla, M. and Lin, C. J. (2026) 'TopK-Guided: Adaptive, Budget-Aware Activation Sparsity for Efficient LLM Inference', Available at: https://omanscience.com/en/articles/topk-guided-adaptive-budget-aware-activation-sparsity-for-efficient-llm-inference.

Vancouver

Agarwalla M, Lin CJ. TopK-Guided: Adaptive, Budget-Aware Activation Sparsity for Efficient LLM Inference. https://omanscience.com/en/articles/topk-guided-adaptive-budget-aware-activation-sparsity-for-efficient-llm-inference

IEEE

M. Agarwalla, and C. J. Lin, "TopK-Guided: Adaptive, Budget-Aware Activation Sparsity for Efficient LLM Inference," https://omanscience.com/en/articles/topk-guided-adaptive-budget-aware-activation-sparsity-for-efficient-llm-inference.