61 citations · 109 across the 4 of their papers we have counts for
5 papers
Random-LTD: Random and Layerwise Token Dropping Brings Efficient Training for Large-scale Transformers
Zhewei Yao, Xiaoxia Wu, Conglong Li +4
Large-scale transformer models have become the de-facto architectures for various machine learning applications, e.g., CV and NLP. However, those large models also introduce prohib…
Bamboo: Making Preemptible Instances Resilient for Affordable Training of Large DNNs
John Thorpe, Pengzhan Zhao, Jonathan Eyolfson +5
DNN models across many domains continue to grow in size, resulting in high resource requirements for effective training, and unpalatable (and often unaffordable) costs for organiza…
ZeRO-Offload: Democratizing Billion-Scale Model Training
Jie Ren, Samyam Rajbhandari, Reza Yazdani Aminabadi +5
Large-scale model training has been a playing ground for a limited few requiring complex model refactoring and access to prohibitively expensive GPU clusters. ZeRO-Offload changes…
Accelerating Training of Transformer-Based Language Models with Progressive Layer Dropping
Minjia Zhang, Yuxiong He
Recently, Transformer-based language models have demonstrated remarkable performance across many NLP domains. However, the unsupervised pre-training step of these models suffers fr…
Navigating with Graph Representations for Fast and Scalable Decoding of Neural Language Models
Minjia Zhang, Xiaodong Liu, Wenhan Wang +2
Neural language models (NLMs) have recently gained a renewed interest by achieving state-of-the-art performance across many natural language processing (NLP) tasks. However, NLMs a…