Abstract
Mixed-precision post-training quantization needs a per-module sensitivity signal; for a text embedder the obvious one -- the retrieval quality a module costs when quantized -- needs relevance labels that deployments rarely have. We measure a label-free substitute: quantization-induced representation drift, obtained by quantizing one module, re-encoding the corpus, and recording how far the output embeddings moved from their full-precision positions. What is specific is the observable: the deployed output representation a dense retriever ranks with. Across five development embedders, configuration-level drift orders sampled mixed-precision plans against held-out retrieval quality at a macro Spearman of 0.911, the sensitivity transports across calibration corpora and retrieval domains in the usable regime, module drifts compose rank-consistently but not numerically, and relevance-derived sensitivity adds no consistent value. The method is one additive allocation under a hard packed-byte budget, with no labels and no search. On three embedders held untouched until method, baselines and hypotheses were frozen and sealed, the pre-registered directional hypothesis against the prior LieQ criterion holds (3/3 at the main budget, no collapse) and drift scores above a two-sided LieQ steelman in 2/3; but at the main budget drift is numerically lower than same-budget uniform precision on all three (-0.99, -0.85, -1.01 points), having reduced module and whole-model drift as designed. Output drift is thus a robust coarse sensitivity signal, not a universally optimal allocation objective: it avoids the catastrophic failures of the transferred signed-geometry adaptation and can remain usable at stressed budgets where uniform collapses, but fine-grained redistribution around a strong uniform operating point remains unresolved.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Han, H., Kim, J., & Jeon, S. H. (2026). Quantize by Drift: Label-Free Mixed-Precision Post-Training Quantization for Text Embedders. https://omanscience.com/en/articles/quantize-by-drift-label-free-mixed-precision-post-training-quantization-for-text-embedders
MLA 9
Han, Hyojung, et al. "Quantize by Drift: Label-Free Mixed-Precision Post-Training Quantization for Text Embedders." https://omanscience.com/en/articles/quantize-by-drift-label-free-mixed-precision-post-training-quantization-for-text-embedders.
Chicago (author–date)
Han, Hyojung, Jongmin Kim, and Seung-Hun Jeon. 2026. "Quantize by Drift: Label-Free Mixed-Precision Post-Training Quantization for Text Embedders." https://omanscience.com/en/articles/quantize-by-drift-label-free-mixed-precision-post-training-quantization-for-text-embedders.
Harvard
Han, H., Kim, J. and Jeon, S. H. (2026) 'Quantize by Drift: Label-Free Mixed-Precision Post-Training Quantization for Text Embedders', Available at: https://omanscience.com/en/articles/quantize-by-drift-label-free-mixed-precision-post-training-quantization-for-text-embedders.
Vancouver
Han H, Kim J, Jeon SH. Quantize by Drift: Label-Free Mixed-Precision Post-Training Quantization for Text Embedders. https://omanscience.com/en/articles/quantize-by-drift-label-free-mixed-precision-post-training-quantization-for-text-embedders
IEEE
H. Han, J. Kim, and S. H. Jeon, "Quantize by Drift: Label-Free Mixed-Precision Post-Training Quantization for Text Embedders," https://omanscience.com/en/articles/quantize-by-drift-label-free-mixed-precision-post-training-quantization-for-text-embedders.