1 paper · 1 filter
Yuchen Xian, Yang He, Yunqiu Xu +1
Speculative decoding (SD) addresses the high inference costs of LLMs by having lightweight drafters generate candidates for large verifiers to validate in parallel. Existing draft-…