1 paper
Anton Plaksin, Sergei Krutikov, Sergei Skvortsov +1
Speculative decoding speeds up autoregressive generation in Large Language Models (LLMs) through a two-step procedure, where a lightweight draft model proposes tokens which the tar…