Preprint Open access
SparLeak: Privacy Leakage from Sparse Attention in LLM Inference on Shared GPUs
Sparse attention is widely used to accelerate long-context inference in modern large language models (LLMs), but its input-dependent execution behavior introduces previously unexplored privacy risks. We identify a new GPU micro-architectural side channel, termed Sparsity-Induced Memory Access (SIMA), which arises from …