6 citations · 8 across the 7 of their papers we have counts for
1 paper · 1 filter
Haifeng Qian, Sujan Kumar Gonugondla, Sungsoo Ha +6
Speculative decoding has emerged as a powerful method to improve latency and throughput in hosting large language models. However, most existing implementations focus on generating…