1 paper
Zixian Li, Tong Li, Chi Xie +2
Speculative decoding accelerates LLM inference only when drafted continuations survive target-model verification. Semi-autoregressive drafters such as DSpark predict an entire toke…