1 paper
Ziyang Ma, Zihong Zhang, Zuchao Li +4
While draft-model-free speculative decoding offers a promising path to efficient LLM inference, it is frequently constrained by stale draft candidates and the high computational co…