Abstract
Conventional Japanese automatic speech recognition (ASR) is supervised by an orthographic transcript, although the same written form can correspond to different lexical readings realized in speech. Such utterances receive an identical target, so their reading distinction is absent from the supervision interface and cannot be recovered reliably by post-hoc text-only grapheme-to-phoneme conversion. We present Ruby-ASR, which refines the conventional target into a span-bound orthographic--lexical-reading sequence. Unlike separate full-sentence orthographic and phonological outputs, the ruby representation locally binds each written span to its realized reading and permits deterministic recovery of both views. We instantiate the target under subtitle-style and verbatim-style transcription conventions using a Qwen3-ASR backbone; a mora-level CTC objective provides auxiliary monotonic reading supervision. The experimental results across five Japanese benchmarks show that refining the recognition target can improve lexical-reading recovery without sacrificing readable orthographic transcription. We release the checkpoints and inference code.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Shi, H., Liu, Y., Yang, X., Liu, J., Hua, C., Chen, X., Liu, L., Zhu, S., & Su, Z. (2026). Ruby-ASR: Evidence-Preserving Supervision for Joint Orthographic and Lexical-Reading Recognition. https://omanscience.com/en/articles/ruby-asr-evidence-preserving-supervision-for-joint-orthographic-and-lexical-reading-recognition
MLA 9
Shi, Hao, et al. "Ruby-ASR: Evidence-Preserving Supervision for Joint Orthographic and Lexical-Reading Recognition." https://omanscience.com/en/articles/ruby-asr-evidence-preserving-supervision-for-joint-orthographic-and-lexical-reading-recognition.
Chicago (author–date)
Shi, Hao, Yun Liu, Xuehao Yang, Jun Liu, Chuanbo Hua, Xuanjun Chen, Lianbo Liu, Shiao Zhu, and Zixiong Su. 2026. "Ruby-ASR: Evidence-Preserving Supervision for Joint Orthographic and Lexical-Reading Recognition." https://omanscience.com/en/articles/ruby-asr-evidence-preserving-supervision-for-joint-orthographic-and-lexical-reading-recognition.
Harvard
Shi, H., Liu, Y., Yang, X., Liu, J., Hua, C., Chen, X., Liu, L., Zhu, S. and Su, Z. (2026) 'Ruby-ASR: Evidence-Preserving Supervision for Joint Orthographic and Lexical-Reading Recognition', Available at: https://omanscience.com/en/articles/ruby-asr-evidence-preserving-supervision-for-joint-orthographic-and-lexical-reading-recognition.
Vancouver
Shi H, Liu Y, Yang X, Liu J, Hua C, Chen X, et al. Ruby-ASR: Evidence-Preserving Supervision for Joint Orthographic and Lexical-Reading Recognition. https://omanscience.com/en/articles/ruby-asr-evidence-preserving-supervision-for-joint-orthographic-and-lexical-reading-recognition
IEEE
H. Shi, Y. Liu, X. Yang, J. Liu, C. Hua, X. Chen, L. Liu, S. Zhu, and Z. Su, "Ruby-ASR: Evidence-Preserving Supervision for Joint Orthographic and Lexical-Reading Recognition," https://omanscience.com/en/articles/ruby-asr-evidence-preserving-supervision-for-joint-orthographic-and-lexical-reading-recognition.