الباحثون

Linxuan Wang

المنشورات 1

نسخة أولية وصول مفتوح

UBTree: Parallel Tree Drafting via Unigram and Bigram Models for Speculative Decoding

Chumeng Liang, Linxuan Wang, Xinyu Peng وآخرون · 2026

Speculative decoding accelerates language model inference by verifying multiple draft tokens in a single target-model pass. Recent parallel drafters have achieved breakthrough performance in frontier production models, but their effectiveness deteriorates as the entropy of target distributions increases due to insuffic …

المؤلفون المشاركون