1 paper
Jiyoung Park, Hankyu Jang, Changseok Song +1
Speculative decoding can substantially accelerate LLM inference, but realizing its benefits in practice is challenging due to evolving workloads and system-level constraints. We pr…