Preprint Open access
CurveTQ: Rotation-Free Trellis Quantization of LLM Weights via Curvature-Weighted Search
The best two-bit weight quantizers for large language models, such as QTIP and Proteus, rotate each weight matrix by a random orthogonal transform, which must be undone at every decoding step, then encode it with a trellis or lattice code under a Euclidean search; the layer Hessian enters only through error feedback be …