1 paper
Amirmohammad Karimi, Chao Gao, Negar Hassanpour
Speculative decoding accelerates large language models' inference by using a lightweight drafter to propose multiple future tokens and a target model to verify them. While recent b…