15 citations · 15 across the 2 of their papers we have counts for
3 papers
SWAN-GPT: An Efficient and Scalable Approach for Long-Context Language Modeling
Krishna C. Puvvada, Faisal Ladhak, Santiago Akle Serrano +8
We present a decoder-only Transformer architecture that robustly generalizes to sequence lengths substantially longer than those seen during training. Our model, SWAN-GPT, interlea…
Stochastic Weight Averaging in Parallel: Large-Batch Training that Generalizes Well
Vipul Gupta, Santiago Akle Serrano, Dennis DeCoste
We propose Stochastic Weight Averaging in Parallel (SWAP), an algorithm to accelerate DNN training. Our algorithm uses large mini-batches to compute an approximate solution quickly…
Democratizing Production-Scale Distributed Deep Learning
Minghuang Ma, Hadi Pouransari, Daniel Chao +5
The interest and demand for training deep neural networks have been experiencing rapid growth, spanning a wide range of applications in both academia and industry. However, trainin…