1 paper · 1 filter
Cunxiao Du, Jing Jiang, Xu Yuanchen +8
Speculative decoding is a relatively new decoding framework that leverages small and efficient draft models to reduce the latency of LLMs. In this study, we introduce GliDe and CaP…