4 papers
CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters
Shaoyuan Huang, Yunfeng Zhao, Na Yan +5
As Large Language Models (LLMs) are increasingly adopted in edge intelligence to power domain-specific applications and personalized services, the quality and efficiency of the LLM…
FineServe: A Fine-Grained Dataset and Characterization of Global LLM Serving Workloads
Tiancheng Zhang, Shaoyuan Huang, Mingyuan Wang +3
Large language models (LLMs) are increasingly deployed as always-on online services, making efficient LLM serving a critical systems challenge. Achieving low latency and high throu…
Stratos: An End-to-End Distillation Pipeline for Customized LLMs under Distributed Cloud Environments
Ziming Dai, Tuo Zhang, Fei Gao +5
The growing industrial demand for customized and cost-efficient large language models (LLMs) is fueled by the rise of vertical, domain-specific tasks and the need to optimize perfo…
HRS: Hybrid Representation Framework with Scheduling Awareness for Time Series Forecasting in Crowdsourced Cloud-Edge Platforms
Tiancheng Zhang, Cheng Zhang, Shuren Liu +3
With the rapid proliferation of streaming services, network load exhibits highly time-varying and bursty behavior, posing serious challenges for maintaining Quality of Service (QoS…