4 papers
Cloud-OpsBench: A Reproducible Benchmark for Agentic Root Cause Analysis in Cloud Systems
Yilun Wang, Guangba Yu, Haiyu Huang +4
LLM agents are increasingly explored for automating root cause analysis (RCA) in cloud-native systems, creating a need to evaluate both diagnostic correctness and the quality of th…
Why Does the LLM Stop Computing: An Empirical Study of User-Reported Failures in Open-Source LLMs
Guangba Yu, Zirui Wang, Yujie Huang +4
The democratization of open-source Large Language Models (LLMs) allows users to fine-tune and deploy models on local infrastructure but exposes them to a First Mile deployment land…
Hierarchical Prediction-based Management for LMaaS Systems
Zhihan Jiang, Yujie Huang, Guangba Yu +3
Large Language Models (LLMs) have revolutionized numerous domains, driving the rise of Language-Model-as-a-Service (LMaaS) platforms that process millions of queries daily. These p…
LLMPrism: Black-box Performance Diagnosis for Production LLM Training Platforms
Zhihan Jiang, Rui Ren, Guangba Yu +8
Large Language Models (LLMs) have brought about revolutionary changes in diverse fields, rendering LLM training of utmost importance for modern enterprises. To meet this demand, mu…