Abstract

Autoregressive text-to-speech (TTS) systems synthesize natural speech but, once trained, offer little control over speaking rate. We show that speaking rate can be steered at inference time, without retraining, by clamping a single decoder block's activation along a discovered speed axis. A decoder-block analysis recovers the rate axis, a neutral operating point, and a per-step intensity scale; at inference, the activation's projection onto this axis is set to a fixed scalar. Learning this direction from synthetically time-stretched and time-compressed speech yields rate control that largely preserves speaker identity, generalizes across model architectures, and maintains high naturalness in objective and human evaluations. Unlike standard additive steering, which breaks at the slow extreme, clamping remains stable on all three systems tested; at moderate targets, the better rule depends on the model. Finally, we show that rate information is decodable across layers but causally steerable only within a mid-depth window, and demonstrate the effectiveness of our approach on the public Seed-TTS-Eval benchmark.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Verdini, F., Asonitis, A., Farhadipour, A., Razavi, M., Honnet, P. E., Avijeet, V., & Gomez, J. P. Z. (2026). Controlling Speaking Rate in Autoregressive TTS via Activation Steering. https://omanscience.com/en/articles/controlling-speaking-rate-in-autoregressive-tts-via-activation-steering

MLA 9

Verdini, Francesco, et al. "Controlling Speaking Rate in Autoregressive TTS via Activation Steering." https://omanscience.com/en/articles/controlling-speaking-rate-in-autoregressive-tts-via-activation-steering.

Chicago (author–date)

Verdini, Francesco, Antonis Asonitis, Aref Farhadipour, Marzieh Razavi, Pierre-Edouard Honnet, Vijeta Avijeet, and Juan Pablo Zuluaga Gomez. 2026. "Controlling Speaking Rate in Autoregressive TTS via Activation Steering." https://omanscience.com/en/articles/controlling-speaking-rate-in-autoregressive-tts-via-activation-steering.

Harvard

Verdini, F., Asonitis, A., Farhadipour, A., Razavi, M., Honnet, P. E., Avijeet, V. and Gomez, J. P. Z. (2026) 'Controlling Speaking Rate in Autoregressive TTS via Activation Steering', Available at: https://omanscience.com/en/articles/controlling-speaking-rate-in-autoregressive-tts-via-activation-steering.

Vancouver

Verdini F, Asonitis A, Farhadipour A, Razavi M, Honnet PE, Avijeet V, et al. Controlling Speaking Rate in Autoregressive TTS via Activation Steering. https://omanscience.com/en/articles/controlling-speaking-rate-in-autoregressive-tts-via-activation-steering

IEEE

F. Verdini, A. Asonitis, A. Farhadipour, M. Razavi, P. E. Honnet, V. Avijeet, and J. P. Z. Gomez, "Controlling Speaking Rate in Autoregressive TTS via Activation Steering," https://omanscience.com/en/articles/controlling-speaking-rate-in-autoregressive-tts-via-activation-steering.