1 paper
Yijiong Yu, Huazheng Wang, Shuai Yuan +2
Speculative Decoding (SD) accelerates low-concurrency LLM inference with a draft-then-verify paradigm. Mainstream methods, however, rely on multi-token prediction, which incurs com…