31 citations · 31 across the 2 of their papers we have counts for
1 paper · 1 filter
Kaixuan Huang, Xudong Guo, Mengdi Wang
Speculative decoding reduces the inference latency of a target large language model via utilizing a smaller and faster draft model. Its performance depends on a hyperparameter K --…