1 paper
Xinyu Wang, Huapeng Zhou, Ziyu Zhao +5
Speculative decoding speeds up generation by letting a cheap draft propose several tokens that a target model checks in one pass. In the single-model form, the draft is a lightweig…