Abstract
Speculative decoding accelerates large language model inference by drafting future tokens cheaply and verifying them with the target model in parallel. Block drafters score a whole block of future tokens in one forward pass, yet standard decoding verifies only the top-scoring chain and discards the other candidates. Because these candidates are already scored, verifying more of them adds target computation but no extra drafting. We introduce CAST (Cost-Aware Speculative Trees), which packs these candidates into a tree and verifies it in a single target pass, leaving the target model, drafter weights, and decoding rule untouched. To decide how wide the tree should be, CAST adds candidates while the expected gain from the next one outweighs the verification time it adds. The width therefore adapts to each deployment from a latency measurement, without sweeping over widths. We evaluate CAST across five domains on three GPU generations and two model families. At its predicted width, CAST is faster than the standard chain in all eight settings, by up to 43%. We also find that the best width depends strongly on the deployment. Where verification cost jumps at a kernel boundary, a 128-token tree is only 2% faster than the standard chain, whereas the tree at the predicted width is 20% faster. Furthermore, we prove that CAST leaves the target output distribution unchanged under both greedy and sampled decoding. Code is available at https://github.com/js-lee-AI/CAST.
Keywords
Subject
Publication details
- Journal
- Not available
- Open access
- Green open access
Cite this article
APA 7
Lee, J., & Eo, S. (2026). CAST: Cost-Aware Speculative Trees from One-Pass Block Drafters. https://omanscience.com/en/articles/cast-cost-aware-speculative-trees-from-one-pass-block-drafters
MLA 9
Lee, Jungseob, and Sugyeong Eo. "CAST: Cost-Aware Speculative Trees from One-Pass Block Drafters." https://omanscience.com/en/articles/cast-cost-aware-speculative-trees-from-one-pass-block-drafters.
Chicago (author–date)
Lee, Jungseob, and Sugyeong Eo. 2026. "CAST: Cost-Aware Speculative Trees from One-Pass Block Drafters." https://omanscience.com/en/articles/cast-cost-aware-speculative-trees-from-one-pass-block-drafters.
Harvard
Lee, J. and Eo, S. (2026) 'CAST: Cost-Aware Speculative Trees from One-Pass Block Drafters', Available at: https://omanscience.com/en/articles/cast-cost-aware-speculative-trees-from-one-pass-block-drafters.
Vancouver
Lee J, Eo S. CAST: Cost-Aware Speculative Trees from One-Pass Block Drafters. https://omanscience.com/en/articles/cast-cost-aware-speculative-trees-from-one-pass-block-drafters
IEEE
J. Lee, and S. Eo, "CAST: Cost-Aware Speculative Trees from One-Pass Block Drafters," https://omanscience.com/en/articles/cast-cost-aware-speculative-trees-from-one-pass-block-drafters.