Preprint Open access
Replay the Curvature: Accurate and Scalable NVFP4 Quantization for Large Language Model Inference
Large language models make weight storage and memory traffic major inference costs, motivating low-precision formats that represent each weight with only a few bits. Such formats use a scale to map floating-point values into a small codebook; NVFP4 improves local range utilization by letting every 16 E2M1 weights share …