2 papers
cs.DC2025
DSDE: Dynamic Speculative Decoding with KLD Stability for Real-World Serving
Mingyu Yang, Jae-Young Choi, Kihyo Moon +2
Speculative decoding accelerates large language model inference, but its reliance on a fixed speculation length is suboptimal in large-batch serving environments with diverse reque…
cs.DC2025
ELIS: Efficient LLM Iterative Scheduling System with Response Length Predictor
Seungbeom Choi, Jeonghoe Goo, Eunjoo Jeon +2
We propose ELIS, a serving system for Large Language Models (LLMs) featuring an Iterative Shortest Remaining Time First (ISRTF) scheduler designed to efficiently manage inference t…