Abstract
Large language models spend the same amount of computation on every token they generate, regardless of how difficult each token is to produce. Methods such as speculative decoding and model routing are built on the premise that much of this computation is unnecessary, yet the computation an individual token actually requires has not been measured. We measure it through a Mixture-of-Agents (MoA) lens: a panel of fifteen language models of increasing capacity, drawn from three families, in which every agent attempts to reproduce a reference sequence token by token, conditioned on the correct preceding tokens. We define the inference cost of the smallest agent that succeeds as the token's sufficient compute, which upper-bounds what the token requires. On three core benchmarks, a 0.5B agent reproduces 92--95\% of reference tokens. Across Qwen, OLMo, and R1-distilled panels, the most expensive 10\% account for 64--80\% of estimated FLOPs. On all 500 MATH-500 problems, the MoA-derived map helps model routing reduce projected latency from 7.59 to 5.12 seconds while slightly improving accuracy, relative to the best confidence-routing baseline. The MoA-map helps drafting use 32.6\% fewer draft tokens and approximately 20\% lower projected latency than fixed-window drafting at similar accuracy. These comparisons reveal remaining allocation headroom, motivating controllers that exploit sufficient-compute structure.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Du, Z., Han, W., Li, H. "., & Chen, Y. (2026). What Does a Token Cost? A Mixture-of-Agents Measurement of Sufficient Per-Token Compute. https://omanscience.com/en/articles/what-does-a-token-cost-a-mixture-of-agents-measurement-of-sufficient-per-token-compute
MLA 9
Du, Zhixu, et al. "What Does a Token Cost? A Mixture-of-Agents Measurement of Sufficient Per-Token Compute." https://omanscience.com/en/articles/what-does-a-token-cost-a-mixture-of-agents-measurement-of-sufficient-per-token-compute.
Chicago (author–date)
Du, Zhixu, Weijia Han, Hai "Helen" Li, and Yiran Chen. 2026. "What Does a Token Cost? A Mixture-of-Agents Measurement of Sufficient Per-Token Compute." https://omanscience.com/en/articles/what-does-a-token-cost-a-mixture-of-agents-measurement-of-sufficient-per-token-compute.
Harvard
Du, Z., Han, W., Li, H. ". and Chen, Y. (2026) 'What Does a Token Cost? A Mixture-of-Agents Measurement of Sufficient Per-Token Compute', Available at: https://omanscience.com/en/articles/what-does-a-token-cost-a-mixture-of-agents-measurement-of-sufficient-per-token-compute.
Vancouver
Du Z, Han W, Li H", Chen Y. What Does a Token Cost? A Mixture-of-Agents Measurement of Sufficient Per-Token Compute. https://omanscience.com/en/articles/what-does-a-token-cost-a-mixture-of-agents-measurement-of-sufficient-per-token-compute
IEEE
Z. Du, W. Han, H. ". Li, and Y. Chen, "What Does a Token Cost? A Mixture-of-Agents Measurement of Sufficient Per-Token Compute," https://omanscience.com/en/articles/what-does-a-token-cost-a-mixture-of-agents-measurement-of-sufficient-per-token-compute.