1 paper · 1 filter
Matan Rusanovsky, Yoav Miron, Roy Uziel +4
Speculative decoding accelerates language-model inference by drafting future tokens that the target model verifies in parallel. A diffusion-style block head such as DFlash is an at…