6 papers
CoLLM: Continuous Adaptation for SLO-Aware LLM Serving on Shared GPU Clusters
Shaoyuan Huang, Yunfeng Zhao, Na Yan +5
As Large Language Models (LLMs) are increasingly adopted in edge intelligence to power domain-specific applications and personalized services, the quality and efficiency of the LLM…
FineServe: A Fine-Grained Dataset and Characterization of Global LLM Serving Workloads
Tiancheng Zhang, Shaoyuan Huang, Mingyuan Wang +3
Large language models (LLMs) are increasingly deployed as always-on online services, making efficient LLM serving a critical systems challenge. Achieving low latency and high throu…
ScaleDL: Towards Scalable and Efficient Runtime Prediction for Distributed Deep Learning Workloads
Xiaokai Wang, Shaoyuan Huang, Yuting Li +1
Deep neural networks (DNNs) form the cornerstone of modern AI services, supporting a wide range of applications, including autonomous driving, chatbots, and recommendation systems.…
HRS: Hybrid Representation Framework with Scheduling Awareness for Time Series Forecasting in Crowdsourced Cloud-Edge Platforms
Tiancheng Zhang, Cheng Zhang, Shuren Liu +3
With the rapid proliferation of streaming services, network load exhibits highly time-varying and bursty behavior, posing serious challenges for maintaining Quality of Service (QoS…
MetaEformer: Unveiling and Leveraging Meta-patterns for Complex and Dynamic Systems Load Forecasting
Shaoyuan Huang, Tiancheng Zhang, Zhongtian Zhang +3
Time series forecasting is a critical and practical problem in many real-world applications, especially for industrial scenarios, where load forecasting underpins the intelligent o…
Sentinel: Scheduling Live Streams with Proactive Anomaly Detection in Crowdsourced Cloud-Edge Platforms
Yuting Li, Shaoyuan Huang, Tengwen Zhang +3
With the rapid growth of live streaming services, Crowdsourced Cloud-edge service Platforms (CCPs) are playing an increasingly important role in meeting the increasing demand. Alth…