507 citations · 694 across the 9 of their papers we have counts for
7 papers · 1 filter
Colossal-Auto: Unified Automation of Parallelization and Activation Checkpoint for Large-scale Models
Yuliang Liu, Shenggui Li, Jiarui Fang +3
In recent years, large-scale models have demonstrated state-of-the-art performance across various domains. However, training such models requires various techniques to address the…
MIGPerf: A Comprehensive Benchmark for Deep Learning Training and Inference Workloads on Multi-Instance GPUs
Huaizheng Zhang, Yuanming Li, Wencong Xiao +7
New architecture GPUs like A100 are now equipped with multi-instance GPU (MIG) technology, which allows the GPU to be partitioned into multiple small, isolated instances. This tech…
EnergonAI: An Inference System for 10-100 Billion Parameter Transformer Models
Jiangsu Du, Ziming Liu, Jiarui Fang +4
Large transformer models display promising performance on a wide range of natural language processing (NLP) tasks. Although the AI community has expanded the model scale to the tri…
Sky Computing: Accelerating Geo-distributed Computing in Federated Learning
Jie Zhu, Shenggui Li, Yang You
Federated learning is proposed by Google to safeguard data privacy through training models locally on users' devices. However, with deep learning models growing in size to achieve…
An Efficient 2D Method for Training Super-Large Deep Learning Models
Qifan Xu, Shenggui Li, Chaoyu Gong +1
Huge neural network models have shown unprecedented performance in real-world applications. However, due to memory constraints, model parallelism must be utilized to host large mod…
Large Batch Optimization for Deep Learning: Training BERT in 76 minutes
Yang You, Jing Li, Sashank Reddi +7
Training large deep neural networks on massive datasets is computationally very challenging. There has been recent surge in interest in using large batch stochastic optimization me…