Preprint Open access
Beyond Masked Sparsity: SNACK Enables Truly Sparse Neural Networks on GPU
Deep neural networks continue to grow in parameter count, driving up training and inference cost on GPUs. Sparse neural networks and Dynamic Sparse Training (DST) promise to reduce these costs, but most implementations rely on binary masks over dense tensors and recover little of the theoretical compute, memory, or ene …