1.3k citations · 3.3k across the 44 of their papers we have counts for
4 papers · 1 filter
Variance-Reduced Gradient Estimation via Noise-Reuse in Online Evolution Strategies
Oscar Li, James Harrison, Jascha Sohl-Dickstein +2
Unrolled computation graphs are prevalent throughout machine learning but present challenges to automatic differentiation (AD) gradient estimation methods when their loss functions…
A Mean Field Theory of Batch Normalization
Greg Yang, Jeffrey Pennington, Vinay Rao +2
We develop a mean field theory for batch normalization in fully-connected feedforward neural networks. In so doing, we provide a precise characterization of signal propagation and…
Understanding and correcting pathologies in the training of learned optimizers
Luke Metz, Niru Maheswaranathan, Jeremy Nixon +2
Deep learning has shown that learned functions can dramatically outperform hand-designed functions on perceptual tasks. Analogously, this suggests that learned optimizers may simil…
Guided evolutionary strategies: Augmenting random search with surrogate gradients
Niru Maheswaranathan, Luke Metz, George Tucker +2
Many applications in machine learning require optimizing a function whose true gradient is unknown, but where surrogate gradient information (directions that may be correlated with…