1 paper
Chendong Sun, Ali Mao, Lei Xu +1
Speculative Decoding is a prominent technique for accelerating the autoregressive inference of large language models (LLMs) by employing a fast draft model to propose candidate tok…