[
    {
        "id": "osp-16323",
        "type": "article-journal",
        "title": "Rounding in Preconditioner Space: Redesigning 4-bit AdamW Optimizer-State Quantization",
        "author": [
            {
                "family": "Li",
                "given": "Hanyang"
            },
            {
                "family": "Tang",
                "given": "Shao"
            },
            {
                "family": "Braithwaite",
                "given": "Daniel Thomas"
            },
            {
                "family": "Dexter",
                "given": "Gregory"
            },
            {
                "family": "Neves",
                "given": "Leonardo"
            },
            {
                "family": "Gupta",
                "given": "Aman"
            },
            {
                "family": "Udagawa",
                "given": "Hiroto"
            },
            {
                "family": "Shivanna",
                "given": "Abhishek"
            },
            {
                "family": "Silva",
                "given": "Daniel"
            },
            {
                "family": "Ramanath",
                "given": "Rohan"
            }
        ],
        "URL": "https://omanscience.com/en/articles/rounding-in-preconditioner-space-redesigning-4-bit-adamw-optimizer-state-quantization",
        "language": "en",
        "issued": {
            "date-parts": [
                [
                    2026
                ]
            ]
        },
        "abstract": "Quantizing AdamW's optimizer states reduces persistent storage, but quantization errors propagate through the moment recurrences and perturb subsequent adaptive updates. We redesign 4-bit optimizer-state quantization for AdamW from the perspective of \\emph{rounding space}: the coordinate in which a quantizer chooses between adjacent reconstruction levels. For the second moment, a local analysis of the quantization cell adjacent to zero shows that small mean state error need not imply small mean preconditioner error at the next step. A one-dimensional quadratic construction further shows qualitatively different optimization dynamics under state-space and preconditioner-space rounding. These results motivate Zero-Inclusive Preconditioner-space Stochastic Rounding (\\textbf{ZIP-SR}), which retains zero in the second-moment codebook and computes stochastic-rounding probabilities in preconditioner space. As a complementary route, Zero-Excluding EDEN calibration (\\textbf{ZE-EDEN}) uses a zero-excluding second-moment codebook and rescales the quantized second-moment block to mitigate the preconditioner distortion caused by the positive quantization floor. Both configurations use 4-bit NormalFloat (NF4) for the first moment, with targeted stochastic rounding of the LM-head first moment during the final 10\\% of training. Across GPT- and Llama-style pretraining experiments ranging from \\textbf{130M} to \\textbf{2.7B} parameters, both methods reduce TorchAO 4-bit AdamW's mean validation-loss gap to 32-bit AdamW at every evaluated model size, with the largest reported gap reduction reaching \\textbf{70\\%}. In full-parameter supervised fine-tuning, both recipes achieve lower validation loss than TorchAO while remaining close to 32-bit AdamW on downstream tasks."
    }
]