26 citations · 31 across the 4 of their papers we have counts for
6 papers · 1 filter
Convex Analysis of the Mean Field Langevin Dynamics
Atsushi Nitanda, Denny Wu, Taiji Suzuki
As an example of the nonlinear Fokker-Planck equation, the mean field Langevin dynamics recently attracts attention due to its connection to (noisy) gradient descent on infinitely…
When Does Preconditioning Help or Hurt Generalization?
Shun-ichi Amari, Jimmy Ba, Roger Grosse +5
While second order optimizers such as natural gradient descent (NGD) often speed up optimization, their effect on generalization has been called into question. This work presents a…
Gradient Descent can Learn Less Over-parameterized Two-layer Neural Networks on Classification Problems
Atsushi Nitanda, Geoffrey Chinot, Taiji Suzuki
Recently, several studies have proven the global convergence and generalization abilities of the gradient descent method for two-layer ReLU networks. Most studies especially focuse…
Functional Gradient Boosting based on Residual Network Perception
Atsushi Nitanda, Taiji Suzuki
Residual Networks (ResNets) have become state-of-the-art models in deep learning and several theoretical studies have been devoted to understanding why ResNet works so well. One at…
Stochastic Particle Gradient Descent for Infinite Ensembles
Atsushi Nitanda, Taiji Suzuki
The superior performance of ensemble methods with infinite models are well known. Most of these methods are based on optimization problems in infinite-dimensional spaces with some…
Accelerated Stochastic Gradient Descent for Minimizing Finite Sums
Atsushi Nitanda
We propose an optimization method for minimizing the finite sums of smooth convex functions. Our method incorporates an accelerated gradient descent (AGD) and a stochastic variance…