1 paper
Yahya Emara, Mauricio Barba da Costa, Chi-Chih Chang +4
Speculative decoding (SD) accelerates language model inference by drafting tokens from a cheap proposal model and verifying them against an expensive target model via rejection sam…