activity
20222024
most citedA Hybrid Tensor-Expert-Data Parallelism Approach to Optimize Mixture-of-Experts Training

39 citations · 94 across the 13 of their papers we have counts for

collaborators

12 papers

cs.LG20231 cited

ZeroQuant-HERO: Hardware-Enhanced Robust Optimized Post-Training Quantization Framework for W8A8 Transformers

Zhewei Yao, Reza Yazdani Aminabadi, Stephen Youn +3

Quantization techniques are pivotal in reducing the memory and computational demands of deep neural network inference. Existing solutions, such as ZeroQuant, offer dynamic quantiza…

cs.AI20238 cited

DeepSpeed4Science Initiative: Enabling Large-Scale Scientific Discovery through Sophisticated AI System Technologies

Shuaiwen Leon Song, Bonnie Kruft, Minjia Zhang +89

In the upcoming decade, deep learning may revolutionize the natural sciences, enhancing our capacity to model and predict natural occurrences. This could herald a new era of scient…

cs.LG20238 cited

DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models

Sam Ade Jacobs, Masahiro Tanaka, Chengming Zhang +4

Computation in a typical Transformer-based large language model (LLM) can be characterized by batch size, hidden dimension, number of layers, and sequence length. Until now, system…

cs.CV202310 cited

RenAIssance: A Survey into AI Text-to-Image Generation in the Era of Large Model

Fengxiang Bie, Yibo Yang, Zhongzhu Zhou +10

Text-to-image generation (TTI) refers to the usage of models that could process text input and generate high fidelity images based on text descriptions. Text-to-image generation us…

cs.LG202310 cited

DeepSpeed-Chat: Easy, Fast and Affordable RLHF Training of ChatGPT-like Models at All Scales

Zhewei Yao, Reza Yazdani Aminabadi, Olatunji Ruwase +16

ChatGPT-like models have revolutionized various applications in artificial intelligence, from summarization and coding to translation, matching or even surpassing human performance…

cs.LG20231 cited

Selective Guidance: Are All the Denoising Steps of Guided Diffusion Important?

Pareesa Ameneh Golnari, Zhewei Yao, Yuxiong He

This study examines the impact of optimizing the Stable Diffusion (SD) guided inference pipeline. We propose optimizing certain denoising steps by limiting the noise computation to…