1 citations · 1 across the 6 of their papers we have counts for
6 papers
Trace Sampling 2.0: Code Knowledge Enhanced Span-level Sampling for Distributed Tracing
Yulun Wu, Guangba Yu, Zhihan Jiang +2
Distributed tracing is an essential diagnostic tool in microservice systems, but the sheer volume of traces places a significant burden on backend storage. A common approach to mit…
KPIRoot+: An Efficient Integrated Framework for Anomaly Detection and Root Cause Analysis in Large-Scale Cloud Systems
Wenwei Gu, Renyi Zhong, Guangba Yu +8
To ensure the reliability of cloud systems, their performance is monitored using KPIs (key performance indicators). When issues arise, root cause localization identifies KPIs respo…
COCA: Generative Root Cause Analysis for Distributed Systems with Code Knowledge
Yichen Li, Yulun Wu, Jinyang Liu +4
Runtime failures are commonplace in modern distributed systems. When such issues arise, users often turn to platforms such as Github or JIRA to report them and request assistance.…
L4: Diagnosing Large-scale LLM Training Failures via Automated Log Analysis
Zhihan Jiang, Junjie Huang, Zhuangbin Chen +6
As Large Language Models (LLMs) show their capabilities across various applications, training customized LLMs has become essential for modern enterprises. However, due to the compl…
FaaSRCA: Full Lifecycle Root Cause Analysis for Serverless Applications
Jin Huang, Pengfei Chen, Guangba Yu +3
Serverless becomes popular as a novel computing paradigms for cloud native services. However, the complexity and dynamic nature of serverless applications present significant chall…
Mint: Cost-Efficient Tracing with All Requests Collection via Commonality and Variability Analysis
Haiyu Huang, Cheng Chen, Kunyi Chen +6
Distributed traces contain valuable information but are often massive in volume, posing a core challenge in tracing framework design: balancing the tradeoff between preserving esse…