Abstract

Speech large language models (speech LLMs) perform well on automatic speech recognition (ASR) when sufficient paired speech-text data is available, but their performance degrades in low-resource settings. A cascaded pipeline that performs speech-to-phoneme (S2P) conversion followed by phoneme-to-grapheme (P2G) conversion has been shown to outperform end-to-end speech LLMs in this regime, suggesting that phoneme-mediated processing is beneficial when paired data is scarce. We propose \textit{phoneme-guided initialization}, a simple method that uses this insight within an end-to-end framework: we pre-train the audio encoder on S2P and the LLM on P2G tasks, then connect them and fine-tune the full model end-to-end on the target ASR task. Experiments on Japanese (CSJ), Chinese (AISHELL-1), and two low-resource languages from Common Voice 25.0 (Tatar and Urdu) show that our method matches or outperforms both the cascaded S2P-P2G baseline and the end-to-end model without P2G initialization.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Magoshi, R., Sakai, S., & Kawahara, T. (2026). Phoneme-Guided Initialization for LLM-based Speech Recognition. https://omanscience.com/en/articles/phoneme-guided-initialization-for-llm-based-speech-recognition

MLA 9

Magoshi, Ryo, et al. "Phoneme-Guided Initialization for LLM-based Speech Recognition." https://omanscience.com/en/articles/phoneme-guided-initialization-for-llm-based-speech-recognition.

Chicago (author–date)

Magoshi, Ryo, Shinsuke Sakai, and Tatsuya Kawahara. 2026. "Phoneme-Guided Initialization for LLM-based Speech Recognition." https://omanscience.com/en/articles/phoneme-guided-initialization-for-llm-based-speech-recognition.

Harvard

Magoshi, R., Sakai, S. and Kawahara, T. (2026) 'Phoneme-Guided Initialization for LLM-based Speech Recognition', Available at: https://omanscience.com/en/articles/phoneme-guided-initialization-for-llm-based-speech-recognition.

Vancouver

Magoshi R, Sakai S, Kawahara T. Phoneme-Guided Initialization for LLM-based Speech Recognition. https://omanscience.com/en/articles/phoneme-guided-initialization-for-llm-based-speech-recognition

IEEE

R. Magoshi, S. Sakai, and T. Kawahara, "Phoneme-Guided Initialization for LLM-based Speech Recognition," https://omanscience.com/en/articles/phoneme-guided-initialization-for-llm-based-speech-recognition.