Abstract
Much of expressive text-to-speech research rests on an untested assumption that written text carries enough information to select an appropriate prosodic style for its delivery. Text-predicted style models improve listener preference, and expressive-appropriateness evaluation presupposes that context constrains style, yet neither measures the assumption itself. This paper tests it as a falsifiable hypothesis against style labels derived from acoustics alone. For each of six speakers in a 1,200-hour conversational corpus, utterances are clustered in the spaces of five speech models, including a prosody-only control, and the cluster of held-out utterances is predicted from twelve text embedding models. Three controls are applied: utterance length is erased from the speech embeddings; accuracy is scored against the majority-class floor of unbalanced clusters rather than uniform chance; and a bag-of-words baseline measures word identity alone. Text predicts the cluster above that floor for all six speakers (+0.111 top-3 accuracy), but bag-of-words achieves three quarters of this. Sentence embeddings add only +0.026, largest for encoders not trained for sentence semantics and reversed by tree-based probes for all others. Acoustic clusters are not compact in text embedding space in any of 360 configurations. The prosody-only space weakens the association for five speakers, but not for the speaker showing it most strongly. Text thus informs these delivery clusters mainly through word choice, whether as a cue to prosody or as a marker of topic and recording situation, and reference-free style selection cannot assume more.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Rehman, A., Zhang, J. J., & Yang, X. (2026). Can Prosodic Style Be Inferred from Text Alone? Evidence from Unsupervised Acoustic Clusters. https://omanscience.com/en/articles/can-prosodic-style-be-inferred-from-text-alone-evidence-from-unsupervised-acoustic-clusters
MLA 9
Rehman, Abdul, et al. "Can Prosodic Style Be Inferred from Text Alone? Evidence from Unsupervised Acoustic Clusters." https://omanscience.com/en/articles/can-prosodic-style-be-inferred-from-text-alone-evidence-from-unsupervised-acoustic-clusters.
Chicago (author–date)
Rehman, Abdul, Jian-Jun Zhang, and Xiaosong Yang. 2026. "Can Prosodic Style Be Inferred from Text Alone? Evidence from Unsupervised Acoustic Clusters." https://omanscience.com/en/articles/can-prosodic-style-be-inferred-from-text-alone-evidence-from-unsupervised-acoustic-clusters.
Harvard
Rehman, A., Zhang, J. J. and Yang, X. (2026) 'Can Prosodic Style Be Inferred from Text Alone? Evidence from Unsupervised Acoustic Clusters', Available at: https://omanscience.com/en/articles/can-prosodic-style-be-inferred-from-text-alone-evidence-from-unsupervised-acoustic-clusters.
Vancouver
Rehman A, Zhang JJ, Yang X. Can Prosodic Style Be Inferred from Text Alone? Evidence from Unsupervised Acoustic Clusters. https://omanscience.com/en/articles/can-prosodic-style-be-inferred-from-text-alone-evidence-from-unsupervised-acoustic-clusters
IEEE
A. Rehman, J. J. Zhang, and X. Yang, "Can Prosodic Style Be Inferred from Text Alone? Evidence from Unsupervised Acoustic Clusters," https://omanscience.com/en/articles/can-prosodic-style-be-inferred-from-text-alone-evidence-from-unsupervised-acoustic-clusters.