Abstract

Tokenizer-free language models remove the inductive bias of fixed tokenizers by modeling text directly as bytes, but the resulting longer sequences substantially increase computation and eliminate explicit text abstractions. We ask whether this additional computation can be useful, and whether standard Transformers can learn the abstractions that tokenization provides. We study these questions on Transformers without specialized tokenization-related architectures. With token-superposition training and hash embeddings, byte Transformers consistently outperform subword Transformers as model size scales. We further find that byte Transformers build local text abstractions as external tokenizers: a set of segmentation-like positions are used to collect local context representations, and restricting up to $25\%$ of intermediate layers to these local representations preserves downstream performance. Finally, these learned structures induce highly non-uniform generation difficulty, with uncertainty concentrated near local structure boundaries; exploiting them for speculative decoding yields $3.4\times$ more accepted tokens than in subword Transformers.

Keywords

Subject

Publication details

Journal
Not available
Open access
Green open access

Cite this article

APA 7

Wang, J., Luo, S., Zhang, Q., & Wu, Y. (2026). Byte Language Models: Scaling, Emergent Abstractions, and Information Allocation. https://omanscience.com/en/articles/byte-language-models-scaling-emergent-abstractions-and-information-allocation

MLA 9

Wang, Jie, et al. "Byte Language Models: Scaling, Emergent Abstractions, and Information Allocation." https://omanscience.com/en/articles/byte-language-models-scaling-emergent-abstractions-and-information-allocation.

Chicago (author–date)

Wang, Jie, Shiwei Luo, Qi Zhang, and Yuanbin Wu. 2026. "Byte Language Models: Scaling, Emergent Abstractions, and Information Allocation." https://omanscience.com/en/articles/byte-language-models-scaling-emergent-abstractions-and-information-allocation.

Harvard

Wang, J., Luo, S., Zhang, Q. and Wu, Y. (2026) 'Byte Language Models: Scaling, Emergent Abstractions, and Information Allocation', Available at: https://omanscience.com/en/articles/byte-language-models-scaling-emergent-abstractions-and-information-allocation.

Vancouver

Wang J, Luo S, Zhang Q, Wu Y. Byte Language Models: Scaling, Emergent Abstractions, and Information Allocation. https://omanscience.com/en/articles/byte-language-models-scaling-emergent-abstractions-and-information-allocation

IEEE

J. Wang, S. Luo, Q. Zhang, and Y. Wu, "Byte Language Models: Scaling, Emergent Abstractions, and Information Allocation," https://omanscience.com/en/articles/byte-language-models-scaling-emergent-abstractions-and-information-allocation.