4 papers
From Positionwise Confidence to Prefix Scheduling: Verifier Skipping in Speculative Decoding
Haoxuan Luo, Jameson Sandler, Ferdinando Fioretto
Speculative decoding is a leading technique to reduce the cost of autoregressive generation by using a small drafter to propose several tokens, which are then verified in parallel…
SpecDiff-2: Scaling Diffusion Drafter Alignment For Faster Speculative Decoding
Jameson Sandler, Jacob K. Christopher, Thomas Hartvigsen +1
Speculative decoding has become the standard approach for accelerating Large Language Model (LLM) inference. It exploits a lossless draft-then-verify procedure to circumvent the la…
The Disparate Impacts of Speculative Decoding
Jameson Sandler, Ahmet Ãstün, Marco Romanelli +2
The practice of speculative decoding, whereby inference is probabilistically supported by a smaller, cheaper, ``drafter'' model, has become a standard technique for systematically…
Beyond Jailbreaking: Auditing Contextual Privacy in LLM Agents
Saswat Das, Jameson Sandler, Ferdinando Fioretto
LLM agents have begun to appear as personal assistants, customer service bots, and clinical aides. While these applications deliver substantial operational benefits, they also requ…