2 papers
cs.DC2026
ELDR: Expert-Locality-Aware Decode Routing for PD-Disaggregated MoE Serving
Sangjin Choi, Sukmin Cho, Yifan Xiong +3
In prefill-decode (PD) disaggregated LLM serving, each request is assigned to a decode worker after prefill. Existing decode routers balance only load; for mixture-of-experts (MoE)…
cs.CL2025
Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding
Sukmin Cho, Sangjin Choi, Taeho Hwang +6
Accelerating inference in Large Language Models (LLMs) is critical for real-time interactions, as they have been widely incorporated into real-world services. Speculative decoding,…