activity
20242026
collaborators

6 papers

cs.DC2026

ELDR: Expert-Locality-Aware Decode Routing for PD-Disaggregated MoE Serving

Sangjin Choi, Sukmin Cho, Yifan Xiong +3

In prefill-decode (PD) disaggregated LLM serving, each request is assigned to a decode worker after prefill. Existing decode routers balance only load; for mixture-of-experts (MoE)…

cs.SE2026

TSGuard: Automated User-Centric Incident Diagnosis for AI Workloads in the Cloud

Yitao Yang, Yangtao Deng, Yifan Xiong +3

AI workloads incur frequent failures and incidents from the underlying infrastructure. The current incident management workflow follows a provider-centric paradigm, where users rep…

cs.CL2025

Sigma-MoE-Tiny Technical Report

Qingguo Hu, Zhenghao Lin, Ziyue Yang +12

Mixture-of-Experts (MoE) has emerged as a promising paradigm for foundation models due to its efficient and powerful scalability. In this work, we present Sigma-MoE-Tiny, an MoE la…

cs.DC2025

SIGMA: An AI-Empowered Training Stack on Early-Life Hardware

Lei Qu, Lianhai Ren, Peng Cheng +12

An increasing variety of AI accelerators is being considered for large-scale training. However, enabling large-scale training on early-life AI accelerators faces three core challen…

cs.LG2025

Argos: Agentic Time-Series Anomaly Detection with Autonomous Rule Generation via Large Language Models

Yile Gu, Yifan Xiong, Jonathan Mace +4

Observability in cloud infrastructure is critical for service providers, driving the widespread adoption of anomaly detection systems for monitoring metrics. However, existing syst…

cs.DC2024

SuperBench: Improving Cloud AI Infrastructure Reliability with Proactive Validation

Yifan Xiong, Yuting Jiang, Ziyue Yang +17

Reliability in cloud AI infrastructure is crucial for cloud service providers, prompting the widespread use of hardware redundancies. However, these redundancies can inadvertently…