18 citations · 18 across the 2 of their papers we have counts for
1 paper · 1 filter
Yuetao Chen, Xuliang Wang, Xinzhou Zheng +3
Speculative decoding has emerged as a pivotal technique to accelerate LLM inference by employing a lightweight draft model to generate candidate tokens that are subsequently verifi…