activity
20242026
most citedPie: Pooling CPU Memory for LLM Inference

1 citations · 1 across the 3 of their papers we have counts for

collaborators

6 papers

cs.DC2026

UCCL-EP: Portable Expert-Parallel Communication

Ziming Mao, Yihan Zhang, Chihan Cui +9

Mixture-of-Experts (MoE) workloads rely on expert parallelism (EP) to achieve high GPU efficiency. State-of-the-art EP communication systems such as DeepEP demonstrate strong perfo…

cs.DC2026

SkyNomad: On Using Multi-Region Spot Instances to Minimize AI Batch Job Cost

Zhifei Li, Tian Xia, Ziming Mao +9

AI batch jobs such as model training, inference pipelines, and data analytics require substantial GPU resources and often need to finish before a deadline. Spot instances offer 3-1…

cs.DC2025

SkyWalker: A Locality-Aware Cross-Region Load Balancer for LLM Inference

Tian Xia, Ziming Mao, Jamison Kerney +5

Serving Large Language Models (LLMs) efficiently in multi-region setups remains a challenge. Due to cost and GPU availability concerns, providers typically deploy LLMs in multiple…

cs.NI2025

An Extensible Software Transport Layer for GPU Networking

Yang Zhou, Zhongjie Chen, Ziming Mao +11

Fast-evolving machine learning (ML) workloads have increasing requirements for networking. However, host network transport on RDMA NICs is hard to evolve, causing problems for ML w…

cs.LG20241 cited

Pie: Pooling CPU Memory for LLM Inference

Yi Xu, Ziming Mao, Xiangxi Mo +2

The rapid growth of LLMs has revolutionized natural language processing and AI analysis, but their increasing size and memory demands present significant challenges. A common solut…

cs.DC2024

SkyServe: Serving AI Models across Regions and Clouds with Spot Instances

Ziming Mao, Tian Xia, Zhanghao Wu +6

Recent years have witnessed an explosive growth of AI models. The high cost of hosting AI services on GPUs and their demanding service requirements, make it timely and challenging…