Authors

JuneHyung Kim

Publications 1

Preprint Open access

Activation Sparsity with Weight Approximation for Faster LLM Decoding on Offloaded Weights

Deploying LLMs on consumer-grade GPUs with insufficient memory to hold their weights can result in prohibitively slow inference, because decoding repeatedly transfers offloaded weights from system RAM or flash storage into GPU at much lower bandwidth than local GPU-memory access. Activation sparsity reduces these trans …

Co-authors