7 citations · 17 across the 17 of their papers we have counts for
1 paper · 1 filter
Jinbin Zhang, Nasib Ullah, Erik Schultheis +1
Speculative decoding accelerates LLM inference by letting a small drafter propose multiple tokens which a large target model verifies once per speculation step. As vocabularies sca…