1 paper
Enyu Zhou, Kai Sheng, Hao Chen +1
Speculative decoding (SD), where a draft model provides multiple candidate tokens for the target model to verify in parallel, has demonstrated significant potential for acceleratin…