2 papers
cs.DC2026
CoCoScale: Leveraging Layer-wise Scaling to Unlock the Potential of Online LLM Serving
Jingfeng Wu, Yiyuan He, Minxian Xu +7
Online large language model (LLM) serving has become the backbone of modern AI applications, powering diverse downstream services through shared hardware clusters. However, modern…
cs.DC2025
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
Yiyuan He, Minxian Xu, Jingfeng Wu +7
Large language models (LLMs) are increasingly deployed in AI infrastructure, driving the need for high throughput, resource efficient serving systems. Disaggregated LLM serving, wh…