439 citations · 917 across the 7 of their papers we have counts for
9 papers
On the Generalization Benefit of Noise in Stochastic Gradient Descent
Samuel L. Smith, Erich Elsen, Soham De
It has long been argued that minibatch stochastic gradient descent can generalize better than large batch gradient descent in deep neural networks. However recent papers have quest…
AlgebraNets
Jordan Hoffmann, Simon Schmitt, Simon Osindero +2
Neural networks have historically been built layerwise from the set of functions in , i.e. with activations and weights/parameters represented…
A Practical Sparse Approximation for Real Time Recurrent Learning
Jacob Menick, Erich Elsen, Utku Evci +3
Current methods for training recurrent neural networks are based on backpropagation through time, which requires storing a complete history of network states, and prohibits updatin…
Sparse GPU Kernels for Deep Learning
Trevor Gale, Matei Zaharia, Cliff Young +1
Scientific workloads have traditionally exploited high levels of sparsity to accelerate computation and reduce memory requirements. While deep neural networks can be made sparse, a…
Fast Sparse ConvNets
Erich Elsen, Marat Dukhan, Trevor Gale +1
Historically, the pursuit of efficient inference has been one of the driving forces behind research into new deep learning architectures and building blocks. Some recent examples i…
High Fidelity Speech Synthesis with Adversarial Networks
Mikołaj Bińkowski, Jeff Donahue, Sander Dieleman +5
Generative adversarial networks have seen rapid development in recent years and have led to remarkable improvements in generative modelling of images. However, their application in…