Abstract

Mixed-precision post-training quantization needs a per-module sensitivity signal; for a text embedder the obvious one -- the retrieval quality a module costs when quantized -- needs relevance labels that deployments rarely have. We measure a label-free substitute: quantization-induced representation drift, obtained by quantizing one module, re-encoding the corpus, and recording how far the output embeddings moved from their full-precision positions. What is specific is the observable: the deployed output representation a dense retriever ranks with. Across five development embedders, configuration-level drift orders sampled mixed-precision plans against held-out retrieval quality at a macro Spearman of 0.911, the sensitivity transports across calibration corpora and retrieval domains in the usable regime, module drifts compose rank-consistently but not numerically, and relevance-derived sensitivity adds no consistent value. The method is one additive allocation under a hard packed-byte budget, with no labels and no search. On three embedders held untouched until method, baselines and hypotheses were frozen and sealed, the pre-registered directional hypothesis against the prior LieQ criterion holds (3/3 at the main budget, no collapse) and drift scores above a two-sided LieQ steelman in 2/3; but at the main budget drift is numerically lower than same-budget uniform precision on all three (-0.99, -0.85, -1.01 points), having reduced module and whole-model drift as designed. Output drift is thus a robust coarse sensitivity signal, not a universally optimal allocation objective: it avoids the catastrophic failures of the transferred signed-geometry adaptation and can remain usable at stressed budgets where uniform collapses, but fine-grained redistribution around a strong uniform operating point remains unresolved.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Han, H., Kim, J., & Jeon, S. H. (2026). Quantize by Drift: Label-Free Mixed-Precision Post-Training Quantization for Text Embedders. https://omanscience.com/en/articles/quantize-by-drift-label-free-mixed-precision-post-training-quantization-for-text-embedders

MLA 9

Han, Hyojung, et al. "Quantize by Drift: Label-Free Mixed-Precision Post-Training Quantization for Text Embedders." https://omanscience.com/en/articles/quantize-by-drift-label-free-mixed-precision-post-training-quantization-for-text-embedders.

Chicago (author–date)

Han, Hyojung, Jongmin Kim, and Seung-Hun Jeon. 2026. "Quantize by Drift: Label-Free Mixed-Precision Post-Training Quantization for Text Embedders." https://omanscience.com/en/articles/quantize-by-drift-label-free-mixed-precision-post-training-quantization-for-text-embedders.

Harvard

Han, H., Kim, J. and Jeon, S. H. (2026) 'Quantize by Drift: Label-Free Mixed-Precision Post-Training Quantization for Text Embedders', Available at: https://omanscience.com/en/articles/quantize-by-drift-label-free-mixed-precision-post-training-quantization-for-text-embedders.

Vancouver

Han H, Kim J, Jeon SH. Quantize by Drift: Label-Free Mixed-Precision Post-Training Quantization for Text Embedders. https://omanscience.com/en/articles/quantize-by-drift-label-free-mixed-precision-post-training-quantization-for-text-embedders

IEEE

H. Han, J. Kim, and S. H. Jeon, "Quantize by Drift: Label-Free Mixed-Precision Post-Training Quantization for Text Embedders," https://omanscience.com/en/articles/quantize-by-drift-label-free-mixed-precision-post-training-quantization-for-text-embedders.