7 papers
CoCoScale: Leveraging Layer-wise Scaling to Unlock the Potential of Online LLM Serving
Jingfeng Wu, Yiyuan He, Minxian Xu +7
Online large language model (LLM) serving has become the backbone of modern AI applications, powering diverse downstream services through shared hardware clusters. However, modern…
Cloud-native and Distributed Systems for Efficient and Scalable Large Language Models -- A Research Agenda
Minxian Xu, Jingfeng Wu, Shengye Song +16
The rapid rise of Large Language Models (LLMs) has revolutionized various artificial intelligence (AI) applications, from natural language processing to code generation. However, t…
Accelerating OpenPangu Inference on NPU via Speculative Decoding
Yuntao Dai, Jing Wu, Hang Gu +1
To mitigate the Memory Wall bottleneck encountered by Large Language Models (LLMs) during inference on \textbf{NPU} hardware, and addressing the scarcity of native support for main…
BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI Infrastructure
Yiyuan He, Minxian Xu, Jingfeng Wu +7
Large language models (LLMs) are increasingly deployed in AI infrastructure, driving the need for high throughput, resource efficient serving systems. Disaggregated LLM serving, wh…
Cloud Native System for LLM Inference Serving
Minxian Xu, Junhan Liao, Jingfeng Wu +3
Large Language Models (LLMs) are revolutionizing numerous industries, but their substantial computational demands create challenges for efficient deployment, particularly in cloud…
Unlock the Potential of Fine-grained LLM Serving via Dynamic Module Scaling
Jingfeng Wu, Yiyuan He, Minxian Xu +3
The rise of large language models (LLMs) has created new opportunities across various fields but has also introduced significant challenges in resource management. Current LLM serv…