11 citations · 11 across the 2 of their papers we have counts for
3 papers
cs.DC2026
LUMEN: Coordinated Failure Recovery for Distributed LLM Serving
Zhang Cao, Shujie Han, Juncheng Zhang +3
Modern large language model (LLM) serving clusters distribute inference requests across multiple worker processes on different GPUs, but failures are prevalent at scale. When a wor…
cs.DC2025
BSODiag: A Global Diagnosis Framework for Batch Servers Outage in Large-scale Cloud Infrastructure Systems
Tao Duan, Runqing Chen, Pinghui Wang +5
Cloud infrastructure is the collective term for all physical devices within cloud systems. Failures within the cloud infrastructure system can severely compromise the stability and…
cs.LG2019★ 11 cited
Robust Data Preprocessing for Machine-Learning-Based Disk Failure Prediction in Cloud Production Environments
Shujie Han, Jun Wu, Erci Xu +7
To provide proactive fault tolerance for modern cloud data centers, extensive studies have proposed machine learning (ML) approaches to predict imminent disk failures for early rem…