1 paper · 1 filter
Zedong Liu, Xinyang Ma, Dejun Luo +9
LLMs are widely adopted in production, pushing inference systems to their limits. Disaggregated LLM serving (e.g., PD separation and KV state disaggregation) improves scalability a…