1 paper
Yuetao Chen, Xuliang Wang, Xinzhou Zheng +3
Speculative decoding has emerged as a pivotal technique to accelerate LLM inference by employing a lightweight draft model to generate candidate tokens that are subsequently verifi…