1 paper
Mingbo Song, Heming Xia, Jun Zhang +4
Speculative Decoding (SD) has emerged as a widely used paradigm to accelerate the inference of large language models (LLMs) without compromising generation quality. It works by eff…