1 paper · 1 filter
Rui Li, Zhaoning Zhang, Libo Zhang +3
Speculative decoding (SD) accelerates LLM inference by verifying draft tokens in parallel. However, this method presents a critical trade-off: it improves throughput in low-load, m…