Preprint Open access
What Does a Token Cost? A Mixture-of-Agents Measurement of Sufficient Per-Token Compute
Large language models spend the same amount of computation on every token they generate, regardless of how difficult each token is to produce. Methods such as speculative decoding and model routing are built on the premise that much of this computation is unnecessary, yet the computation an individual token actually re …