171 citations · 237 across the 13 of their papers we have counts for
1 paper · 1 filter
Xiaoxuan Liu, Lanxiang Hu, Peter Bailis +4
Speculative decoding is a pivotal technique to accelerate the inference of large language models (LLMs) by employing a smaller draft model to predict the target model's outputs. Ho…