Abstract

Text-output scores alone do not show whether quantization preserves performance on speech tasks whose target labels cannot be recovered from the transcript. We evaluate fixed mixed 4/8-bit Qwen2-Audio-7B-Instruct allocations averaging 6 and 7 bits per parameter on 508 English-to-German FLEURS utterances and on 512 RAVDESS emotion clips from 16 speakers. The BLEU and chrF differences from half precision (FP16) have intervals that include zero for both allocations. On RAVDESS, the same two sentences occur equally often with every emotion label. The absolute accuracy differences from FP16 are -3.71% for 6 bit and -1.17% for 7 bit. The 6-bit speaker interval excludes zero and an exact two-sided sign-flip test gives p=0.0148; the 7-bit interval includes zero. Same-budget controls do not identify either selected allocation as best. This case study shows why translation scores and performance on tasks beyond the transcript need separate evaluation.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Geng, M., Ji, J., & Xu, J. (2026). Text Scores Do Not Establish Performance on Lexically Non-Diagnostic Speech Tasks: A Qwen2-Audio Quantization Case Study. https://omanscience.com/en/articles/text-scores-do-not-establish-performance-on-lexically-non-diagnostic-speech-tasks-a-qwen2-audio-quantization-case-study

MLA 9

Geng, Mengzhe, et al. "Text Scores Do Not Establish Performance on Lexically Non-Diagnostic Speech Tasks: A Qwen2-Audio Quantization Case Study." https://omanscience.com/en/articles/text-scores-do-not-establish-performance-on-lexically-non-diagnostic-speech-tasks-a-qwen2-audio-quantization-case-study.

Chicago (author–date)

Geng, Mengzhe, Jinxi Ji, and Junhao Xu. 2026. "Text Scores Do Not Establish Performance on Lexically Non-Diagnostic Speech Tasks: A Qwen2-Audio Quantization Case Study." https://omanscience.com/en/articles/text-scores-do-not-establish-performance-on-lexically-non-diagnostic-speech-tasks-a-qwen2-audio-quantization-case-study.

Harvard

Geng, M., Ji, J. and Xu, J. (2026) 'Text Scores Do Not Establish Performance on Lexically Non-Diagnostic Speech Tasks: A Qwen2-Audio Quantization Case Study', Available at: https://omanscience.com/en/articles/text-scores-do-not-establish-performance-on-lexically-non-diagnostic-speech-tasks-a-qwen2-audio-quantization-case-study.

Vancouver

Geng M, Ji J, Xu J. Text Scores Do Not Establish Performance on Lexically Non-Diagnostic Speech Tasks: A Qwen2-Audio Quantization Case Study. https://omanscience.com/en/articles/text-scores-do-not-establish-performance-on-lexically-non-diagnostic-speech-tasks-a-qwen2-audio-quantization-case-study

IEEE

M. Geng, J. Ji, and J. Xu, "Text Scores Do Not Establish Performance on Lexically Non-Diagnostic Speech Tasks: A Qwen2-Audio Quantization Case Study," https://omanscience.com/en/articles/text-scores-do-not-establish-performance-on-lexically-non-diagnostic-speech-tasks-a-qwen2-audio-quantization-case-study.