1 paper
Feiye Huo, Jianchao Tan, Kefeng Zhang +2
The growing scale of Large Language Models (LLMs) has exacerbated inference latency and computational costs. Speculative decoding methods, which aim to mitigate these issues, often…