Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
TreeFlash: Parallel AR-Approximation for Faster Speculative Decoding
Peer Rheinboldt, Frédéric Berdoz, Roger Wattenhofer
One-shot block drafters for speculative decoding generate the full draft in a single forward pass, achieving strong throughput by eliminating sequential token generation. However,…
cs.LG2025
Steering Pretrained Drafters during Speculative Decoding
Frédéric Berdoz, Peer Rheinboldt, Roger Wattenhofer
Speculative decoding accelerates language model inference by separating generation into fast drafting and parallel verification. Its main limitation is drafter-verifier misalignmen…