6 citations · 11 across the 9 of their papers we have counts for
1 paper · 2 filters
Haifeng Qian, Sujan Kumar Gonugondla, Sungsoo Ha +6
Speculative decoding has emerged as a powerful method to improve latency and throughput in hosting large language models. However, most existing implementations focus on generating…