1 paper
Rahul Krishna Thomas, Arka Pal
Speculative sampling reduces the latency of autoregressive decoding for target model LLMs without sacrificing inference quality, by using a cheap draft model to suggest a candidate…