2 papers
cs.LG2026
Semantic Parallelism: Redefining Efficient MoE Inference via Model-Data Co-Scheduling
Yan Li, Zhenyu Zhang, Zhengang Wang +2
Prevailing LLM serving engines employ expert parallelism (EP) to implement multi-device inference of massive MoE models. However, the efficiency of expert parallel inference is lar…
cs.SE2026
A Survey on Failure Analysis and Fault Injection in AI Systems
Guangba Yu, Gou Tan, Haojia Huang +4
The rapid advancement of Artificial Intelligence (AI) has led to its integration into various areas, especially with Large Language Models (LLMs) significantly enhancing capabiliti…