4 citations · 5 across the 3 of their papers we have counts for
1 paper · 1 filter
Situo Zhang, Hankun Wang, Da Ma +4
Speculative Decoding (SD) is a popular lossless technique for accelerating the inference of Large Language Models (LLMs). We show that the decoding speed of SD frameworks with stat…