collaborators

5 papers

cs.DC2026

CoCoScale: Leveraging Layer-wise Scaling to Unlock the Potential of Online LLM Serving

Jingfeng Wu, Yiyuan He, Minxian Xu +7

Online large language model (LLM) serving has become the backbone of modern AI applications, powering diverse downstream services through shared hardware clusters. However, modern…

cs.DC2025

BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure

Yiyuan He, Minxian Xu, Jingfeng Wu +7

Large language models (LLMs) are increasingly deployed in AI infrastructure, driving the need for high throughput, resource efficient serving systems. Disaggregated LLM serving, wh…

cs.DC2025

Cloud Native System for LLM Inference Serving

Minxian Xu, Junhan Liao, Jingfeng Wu +3

Large Language Models (LLMs) are revolutionizing numerous industries, but their substantial computational demands create challenges for efficient deployment, particularly in cloud…

cs.DC2025

Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling

Jingfeng Wu, Yiyuan He, Minxian Xu +3

The rise of large language models (LLMs) has created new opportunities across various fields but has also introduced significant challenges in resource management. Current LLM serv…

cs.DC2025

CloudNativeSim: a toolkit for modeling and simulation of cloud-native applications

Jingfeng Wu, Minxian Xu, Yiyuan He +2

Cloud-native applications are increasingly becoming popular in modern software design. Employing a microservice-based architecture into these applications is a prevalent strategy t…