Preprint Open access
Activation Sparsity with Weight Approximation for Faster LLM Decoding on Offloaded Weights
Deploying LLMs on consumer-grade GPUs with insufficient memory to hold their weights can result in prohibitively slow inference, because decoding repeatedly transfers offloaded weights from system RAM or flash storage into GPU at much lower bandwidth than local GPU-memory access. Activation sparsity reduces these trans …