1 paper
Wei Zhong, Manasa Bharadwaj, Yixiao Wang +2
Speculative decoding (SD) is a widely adopted approach for accelerating inference in large language models (LLMs), particularly when the draft and target models are well aligned. H…