1 paper · 1 filter
Chendong Sun, Ali Mao, Lei Xu +1
Speculative Decoding is a prominent technique for accelerating the autoregressive inference of large language models (LLMs) by employing a fast draft model to propose candidate tok…