2 papers
cs.AI2026
SEIS: Self-Evolving Inference Systems
Zhen Xu, Jingyu Liu, Zongze Li +2
Inference systems determine how fast and how cheaply language models can be served, so making them faster has direct practical value. However, prior work focuses mostly on optimizi…
cs.NI2026
Not All Prefills Are Equal: PPD Disaggregation for Multi-turn LLM Serving
Zongze Li, Jingyu Liu, Zhen Xu +3
Prefill-Decode (PD) disaggregation has become the standard architecture for modern LLM inference engines, which alleviates the interference of two distinctive workloads. With the g…