Preprint Open access
Weight Decay and Neuron Condensation: A Three-Stage Analysis of Two-Layer ReLU Networks
Weight decay is widely used as a regularization technique in neural network training, yet its role in neuron condensation (parameter direction alignment) remains unclear. Starting from a parameter initialization in the neural tangent kernel regime, we characterize training dynamics under weight decay through three stag …