Abstract

Post-training quantization to the GGUF format's mixed-precision K-quants is commonly how open-weight language models reach consumer hardware, yet its effect on fine-grained lexical competence is uncharacterized. We audit 27 quantized artifacts across 13 families and four architecture backbones, 0.35B-14B parameters, evaluated down their published ladder to Q2_K (about 2.6 bits per weight), on 429 frequency-validated rare English words under two probes: surface inclusion of a prompt-supplied word and its one-sentence definition, scored by a tiered multi-synonym matcher, its error measured by a blind LLM-judge census of every definition, with human verification. Three regimes emerge at Q2: total collapse into unusable builds, severe semantic dissociation in sub-2B models, and mostly robust preservation above about 3B. In every sub-2B artifact, definitions fall 20-67% below the artifact's baseline, typically several times the inclusion loss. Two controls separate rarity from task difficulty: within the rare set, loss rises with rarity in six of seven sub-2B artifacts, and on a 100-word common-word set rare words lose more than common words in all eight, significantly in six. Tokenizer vocabulary size does not predict the damage (Spearman rho=0.12); parameter count dominates (rho=0.72), confirmed within five of six same-tokenizer families. Q4_K_M remains lexically clean at >=1B. The damage is frequency-graded, provider-dependent, and not calibrated by WikiText-2 perplexity: across nine artifact-matched ladders, near-identical Q2 penalties (44.7%/47.6%) separate an artifact keeping its definitions (3.6%) from one losing them (43.6%). Aggressively quantized small models can keep generating fluent text while no longer knowing what it means, risking hardware-constrained deployments in domains where semantics carries consequences. Validation must be per artifact.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Singh, S. K., Dadwhal, Y. S., & Vedak, M. (2026). Saying, Not Knowing: Aggressively GGUF-Quantized Small Language Models Still Write Rare Words They Can No Longer Define. https://omanscience.com/en/articles/saying-not-knowing-aggressively-gguf-quantized-small-language-models-still-write-rare-words-they-can-no-longer-define

MLA 9

Singh, Saurabh Kumar, et al. "Saying, Not Knowing: Aggressively GGUF-Quantized Small Language Models Still Write Rare Words They Can No Longer Define." https://omanscience.com/en/articles/saying-not-knowing-aggressively-gguf-quantized-small-language-models-still-write-rare-words-they-can-no-longer-define.

Chicago (author–date)

Singh, Saurabh Kumar, Yogeshwar Singh Dadwhal, and Malhar Vedak. 2026. "Saying, Not Knowing: Aggressively GGUF-Quantized Small Language Models Still Write Rare Words They Can No Longer Define." https://omanscience.com/en/articles/saying-not-knowing-aggressively-gguf-quantized-small-language-models-still-write-rare-words-they-can-no-longer-define.

Harvard

Singh, S. K., Dadwhal, Y. S. and Vedak, M. (2026) 'Saying, Not Knowing: Aggressively GGUF-Quantized Small Language Models Still Write Rare Words They Can No Longer Define', Available at: https://omanscience.com/en/articles/saying-not-knowing-aggressively-gguf-quantized-small-language-models-still-write-rare-words-they-can-no-longer-define.

Vancouver

Singh SK, Dadwhal YS, Vedak M. Saying, Not Knowing: Aggressively GGUF-Quantized Small Language Models Still Write Rare Words They Can No Longer Define. https://omanscience.com/en/articles/saying-not-knowing-aggressively-gguf-quantized-small-language-models-still-write-rare-words-they-can-no-longer-define

IEEE

S. K. Singh, Y. S. Dadwhal, and M. Vedak, "Saying, Not Knowing: Aggressively GGUF-Quantized Small Language Models Still Write Rare Words They Can No Longer Define," https://omanscience.com/en/articles/saying-not-knowing-aggressively-gguf-quantized-small-language-models-still-write-rare-words-they-can-no-longer-define.