6 citations · 6 across the 4 of their papers we have counts for
1 paper · 1 filter
Weijie Shi, Qiang Xu, Fan Deng +9
Speculative decoding accelerates LLM inference by drafting a tree of candidate continuations and verifying it in one target forward. Existing drafters fall into two camps with oppo…