papers
Publications (3)
cs.LG2026
Enabling Performant and Flexible Model-Internal Observability for LLM Inference
Nengneng Yu, Sixian Xiong, Yibo Zhao +2
Today's inference-time workloads increasingly depend on timely access to a model's internal states. We present DMI-Lib, a high-speed deep model inspector that treats internal obser…
cs.DC2026
Don't Let a Few Network Failures Slow the Entire AllReduce
Peiqing Chen, Jiedong Jiang, Nengneng Yu +4
Network failures are among the most frequent hardware faults in large-scale GPU clusters and a leading cause of training-job interruptions. Modern collective communication librarie…
cs.DC2025
Reliable and Resilient Collective Communication Library for LLM Training and Serving
Wei Wang, Nengneng Yu, Sixian Xiong +1
Modern ML training and inference now span tens to tens of thousands of GPUs, where network faults can waste 10--15\% of GPU hours due to slow recovery. Common network errors and li…