1 paper
Zhi-Kai Chen, Jun-Jie Tao, Wei-Xiang Mao +2
The efficiency of Large Language Model (LLM) serving is fundamentally limited by the sequential nature of autoregressive decoding. Speculative Decoding (SD) mitigates this by using…