1 paper · 1 filter
Evangelos Georganas, Dhiraj Kalamkar, Alexander Kozlov +1
Speculative decoding (SD) has emerged as a method to accelerate LLM inference without sacrificing any accuracy over the 16-bit model inference. In a typical SD setup, the idea is t…