1 paper
Yuan Lyu, Bharath Irukulapati, Jaya Prakash Champati
Speculative decoding (SD) accelerates LLM inference by 1.5-3 times when the draft and target models are co-located. This has motivated a distributed variant (DSD) that places t…