1 citations · 1 across the 1 of their papers we have counts for
1 paper · 1 filter
Clara Mohri, Haim Kaplan, Tal Schuster +2
Transformer language models generate text autoregressively, making inference latency proportional to the number of tokens generated. Speculative decoding reduces this latency witho…