4 citations · 5 across the 7 of their papers we have counts for
5 papers · 1 filter
UCCL-Zip: Lossless Compression Supercharged GPU Communication
Shuang Ma, Chon Lam Lao, Zhiying Xu +8
The rapid growth of large language models (LLMs) has made GPU communication a critical bottleneck. While prior work reduces communication volume via quantization or lossy compressi…
MXDAG: A Hybrid Abstraction for Cluster Applications
Weitao Wang, Sushovan Das, Xinyu Crystal Wu +3
Distributed applications, such as database queries and distributed training, consist of both compute and network tasks. DAG-based abstraction primarily targets compute tasks and ha…
Efficient and Less Centralized Federated Learning
Li Chou, Zichang Liu, Zhuang Wang +1
With the rapid growth in mobile computing, massive amounts of data and computing resources are now located at the edge. To this end, Federated learning (FL) is becoming a widely ad…
MergeComp: A Compression Scheduler for Scalable Communication-Efficient Distributed Training
Zhuang Wang, Xinyu Wu, T. S. Eugene Ng
Large-scale distributed training is increasingly becoming communication bound. Many gradient compression algorithms have been proposed to reduce the communication overhead and impr…
Fair Packet Scheduling in Network on Chip
Zhuang Wang, Xiao Lv, Mingyu Yan +2
Interconnection networks of parallel systems are used for servicing traf- fic generated by different applications, often belonging to different users. When multiple traffic flows c…