15 citations · 26 across the 3 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2021★ 1 cited
Training With Data Dependent Dynamic Learning Rates
Shreyas Saxena, Nidhi Vyas, Dennis DeCoste
Recently many first and second order variants of SGD have been proposed to facilitate training of Deep Neural Networks (DNNs). A common limitation of these works stem from the fact…
cs.LG2020★ 15 cited
Stochastic Weight Averaging in Parallel: Large-Batch Training that Generalizes Well
Vipul Gupta, Santiago Akle Serrano, Dennis DeCoste
We propose Stochastic Weight Averaging in Parallel (SWAP), an algorithm to accelerate DNN training. Our algorithm uses large mini-batches to compute an approximate solution quickly…