1 paper
Kaiqi Zhang, Jing Zhao, Rui Chen
Large Language Models (LLMs) exhibit high inference latency due to their autoregressive decoding nature. While the draft head in speculative decoding mitigates this issue, its full…