1 paper
Yuesong Liu, Yuan Zeng, Min Lyu +3
Speculative decoding alleviates the memory-bandwidth bottleneck in large language model inference, but its acceleration is jointly constrained by drafting overhead, token acceptanc…