3 citations · 3 across the 1 of their papers we have counts for
4 papers
Adaptive Gradient Methods Converge Faster with Over-Parameterization (but you should do a line-search)
Sharan Vaswani, Issam Laradji, Frederik Kunstner +3
Adaptive gradient methods are typically used for training over-parameterized models. To better understand their behaviour, we study a simplistic setting -- smooth, convex losses wi…
BackPACK: Packing more into backprop
Felix Dangel, Frederik Kunstner, Philipp Hennig
Automatic differentiation frameworks are optimized for exactly one thing: computing the average mini-batch gradient. Yet, other quantities such as the variance of the mini-batch gr…
Limitations of the Empirical Fisher Approximation for Natural Gradient Descent
Frederik Kunstner, Lukas Balles, Philipp Hennig
Natural gradient descent, which preconditions a gradient descent update with the Fisher information matrix of the underlying statistical model, is a way to capture partial second-o…
SLANG: Fast Structured Covariance Approximations for Bayesian Deep Learning with Natural Gradient
Aaron Mishkin, Frederik Kunstner, Didrik Nielsen +2
Uncertainty estimation in large deep-learning models is a computationally challenging task, where it is difficult to form even a Gaussian approximation to the posterior distributio…