activity
20172024
most citedA Hybrid Tensor-Expert-Data Parallelism Approach to Optimize Mixture-of-Experts Training

39 citations · 109 across the 8 of their papers we have counts for

collaborators
Showing 2023Show all

5 papers · 1 filter

cs.AI20238 cited

DeepSpeed4Science Initiative: Enabling Large-Scale Scientific Discovery through Sophisticated AI System Technologies

Shuaiwen Leon Song, Bonnie Kruft, Minjia Zhang +89

In the upcoming decade, deep learning may revolutionize the natural sciences, enhancing our capacity to model and predict natural occurrences. This could herald a new era of scient…

cs.CV2023

DeepSpeed-VisualChat: Multi-Round Multi-Image Interleave Chat via Multi-Modal Causal Attention

Zhewei Yao, Xiaoxia Wu, Conglong Li +6

Most of the existing multi-modal models, hindered by their incapacity to adeptly manage interleaved image-and-text inputs in multi-image, multi-round dialogues, face substantial co…

cs.LG202310 cited

DeepSpeed-Chat: Easy, Fast and Affordable RLHF Training of ChatGPT-like Models at All Scales

Zhewei Yao, Reza Yazdani Aminabadi, Olatunji Ruwase +16

ChatGPT-like models have revolutionized various applications in artificial intelligence, from summarization and coding to translation, matching or even surpassing human performance…

cs.DC2023

MCR-DL: Mix-and-Match Communication Runtime for Deep Learning

Quentin Anthony, Ammar Ahmad Awan, Jeff Rasley +5

In recent years, the training requirements of many state-of-the-art Deep Learning (DL) models have scaled beyond the compute and memory capabilities of a single processor, and nece…

cs.LG202339 cited

A Hybrid Tensor-Expert-Data Parallelism Approach to Optimize Mixture-of-Experts Training

Siddharth Singh, Olatunji Ruwase, Ammar Ahmad Awan +3

Mixture-of-Experts (MoE) is a neural network architecture that adds sparsely activated expert blocks to a base model, increasing the number of parameters without impacting computat…