Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
SpecDiff-2: Scaling Diffusion Drafter Alignment For Faster Speculative Decoding
Jameson Sandler, Jacob K. Christopher, Thomas Hartvigsen +1
Speculative decoding has become the standard approach for accelerating Large Language Model (LLM) inference. It exploits a lossless draft-then-verify procedure to circumvent the la…
cs.CL2025
The Disparate Impacts of Speculative Decoding
Jameson Sandler, Ahmet Üstün, Marco Romanelli +2
The practice of speculative decoding, whereby inference is probabilistically supported by a smaller, cheaper, ``drafter'' model, has become a standard technique for systematically…