1 paper
Yijun Lin, Jinhao Sheng, Qingyue Cai +1
Autoregressive language models suffer from high inference latency due to their sequential decoding nature. Speculative decoding (SD) mitigates this by employing a lightweight draft…