1 paper
Longze Chen, Renke Shan, Huiming Wang +6
Speculative decoding (SD) is a promising method for accelerating the decoding process of Large Language Models (LLMs). The efficiency of SD primarily hinges on the consistency betw…