Abstract

German is a morphologically rich language whose syllable structure is exceptionally well-predicted by the Knuth--Liang hyphenation algorithm. We ask whether phonologically informed tokenization can serve as a competitive target for end-to-end speech recognition. We compare three tokenizer families on the Omnilingual ASR wav2vec 2.0 backbone fine-tuned with CTC: the pretrained multilingual character inventory, a data-driven Byte-Pair Encoding (BPE) over orthography, and phonologically informed units from Pyphen syllabification and grapheme-to-phoneme conversion. Across 40 fine-tunes, we evaluate on three German test sets spanning orthogonal shifts: in-domain read speech, dialectal spontaneous speech, and standard-German spontaneous speech. In-domain, all phonologically informed tokenizers match BPE and the multilingual character baseline on both WER and CER. Under domain shift the picture splits along vocabulary size rather than the linguistic axis of variation: at small vocabularies, syllable-aware tokenization improves on dialectal speech, where phonetic surface forms vary but syllable structure is preserved, and stays ahead on spontaneous speech, where new word-forms violate vocabulary closure. A phoneme-level confusion analysis further shows that all tokenizers commit the same canonical function-word errors, indicating that the acoustic encoder, not the tokenizer, dominates the error topology. Our findings suggest that tokenizer choice may depend on the vocabulary budget as much as on the distribution shift expected at deployment rather than reducing to a single universal optimum.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Witzl, C., Bocklet, T., & Riedhammer, K. (2026). Phonologically Informed Tokenization for German Speech Recognition: A Cross-Domain Study. https://omanscience.com/en/articles/phonologically-informed-tokenization-for-german-speech-recognition-a-cross-domain-study

MLA 9

Witzl, Christopher, et al. "Phonologically Informed Tokenization for German Speech Recognition: A Cross-Domain Study." https://omanscience.com/en/articles/phonologically-informed-tokenization-for-german-speech-recognition-a-cross-domain-study.

Chicago (author–date)

Witzl, Christopher, Tobias Bocklet, and Korbinian Riedhammer. 2026. "Phonologically Informed Tokenization for German Speech Recognition: A Cross-Domain Study." https://omanscience.com/en/articles/phonologically-informed-tokenization-for-german-speech-recognition-a-cross-domain-study.

Harvard

Witzl, C., Bocklet, T. and Riedhammer, K. (2026) 'Phonologically Informed Tokenization for German Speech Recognition: A Cross-Domain Study', Available at: https://omanscience.com/en/articles/phonologically-informed-tokenization-for-german-speech-recognition-a-cross-domain-study.

Vancouver

Witzl C, Bocklet T, Riedhammer K. Phonologically Informed Tokenization for German Speech Recognition: A Cross-Domain Study. https://omanscience.com/en/articles/phonologically-informed-tokenization-for-german-speech-recognition-a-cross-domain-study

IEEE

C. Witzl, T. Bocklet, and K. Riedhammer, "Phonologically Informed Tokenization for German Speech Recognition: A Cross-Domain Study," https://omanscience.com/en/articles/phonologically-informed-tokenization-for-german-speech-recognition-a-cross-domain-study.