Authors

Weiqiao Shan

Publications 2

Preprint Open access

Rethinking Length-Based Training: Batch Composition and Loss Normalization in Speech Token Language Models

Short-to-long training is a simple curriculum for speech models, but its gains can be difficult to interpret. In speech token language models, length-based training can change the shuffle policy, batch composition, token retention, and token weights under batch-mean loss. We disentangle these factors through matched co …

Co-authors