الباحثون

Sergio Maffeis

المنشورات 1

نسخة أولية وصول مفتوح

Speedbumps: Rejection Attacks on Speculative Decoding

Speculative decoding is a popular technique for increasing the speed and reducing the costs of large language model (LLM) inference by verifying multiple draft tokens in a single target-model forward pass. The resulting benefit depends on the ability of the drafter to approximate the target model's distribution. In thi …

المؤلفون المشاركون