1 paper
Xiaofan Lu, Yixiao Zeng, Feiyang Ma +2
Speculative Decoding (SD) is a technique to accelerate the inference of Large Language Models (LLMs) by using a lower complexity draft model to propose candidate tokens verified by…