1 paper · 1 filter
Zilin Xiao, Hongming Zhang, Tao Ge +3
Speculative decoding has proven to be an efficient solution to large language model (LLM) inference, where the small drafter predicts future tokens at a low cost, and the target mo…