1 citations · 1 across the 2 of their papers we have counts for
1 paper · 1 filter
Sukmin Cho, Sangjin Choi, Taeho Hwang +6
Accelerating inference in Large Language Models (LLMs) is critical for real-time interactions, as they have been widely incorporated into real-world services. Speculative decoding,…