نسخة أولية وصول مفتوح
OSFP4: Joint Optimization of Diagonal Smoothing and Block Scales for NVFP4 Quantization
NVFP4 is an attractive datatype for large language model (LLM) inference, offering compact storage and native tensor-core acceleration. However, preserving accuracy using NVFP4 requires careful quantization. In this work we develop a novel quantization scheme called Optimized Smoothing and Scaling for NVFP4 (OSFP4). Fo …