1 paper
Saw S. Lin, Jyh-Shing Roger Jang
Speculative decoding accelerates LLM inference by drafting tokens and verifying them in parallel. Block-diffusion drafters such as DFlash model only per-position marginals, and tre…