471 citations · 1k across the 22 of their papers we have counts for
3 papers · 1 filter
Differentiable Annealed Importance Sampling and the Perils of Gradient Noise
Guodong Zhang, Kyle Hsu, Jianing Li +2
Annealed importance sampling (AIS) and related algorithms are highly effective tools for marginal likelihood estimation, but are not fully differentiable due to the use of Metropol…
When Does Preconditioning Help or Hurt Generalization?
Shun-ichi Amari, Jimmy Ba, Roger Grosse +5
While second order optimizers such as natural gradient descent (NGD) often speed up optimization, their effect on generalization has been called into question. This work presents a…
Fast Convergence of Natural Gradient Descent for Overparameterized Neural Networks
Guodong Zhang, James Martens, Roger Grosse
Natural gradient descent has proven effective at mitigating the effects of pathological curvature in neural network optimization, but little is known theoretically about its conver…