11 citations · 14 across the 4 of their papers we have counts for
6 papers · 1 filter
High-dimensional Asymptotics of Feature Learning: How One Gradient Step Improves the Representation
Jimmy Ba, Murat A. Erdogdu, Taiji Suzuki +3
We study the first gradient descent step on the first-layer parameters in a two-layer neural network: $f(\boldsymbol{x}) = \frac{1}{\sqrt{N}}\boldsymbol{a}^\topσ(\…
Convex Analysis of the Mean Field Langevin Dynamics
Atsushi Nitanda, Denny Wu, Taiji Suzuki
As an example of the nonlinear Fokker-Planck equation, the mean field Langevin dynamics recently attracts attention due to its connection to (noisy) gradient descent on infinitely…
When Does Preconditioning Help or Hurt Generalization?
Shun-ichi Amari, Jimmy Ba, Roger Grosse +5
While second order optimizers such as natural gradient descent (NGD) often speed up optimization, their effect on generalization has been called into question. This work presents a…
On the Optimal Weighted Regularization in Overparameterized Linear Regression
Denny Wu, Ji Xu
We consider the linear model with in the overparameterized regime . We estimate $\ma…
Stochastic Runge-Kutta Accelerates Langevin Monte Carlo and Beyond
Xuechen Li, Denny Wu, Lester Mackey +1
Sampling with Markov chain Monte Carlo methods often amounts to discretizing some continuous-time dynamics with numerical integration. In this paper, we establish the convergence r…
Post Selection Inference with Incomplete Maximum Mean Discrepancy Estimator
Makoto Yamada, Denny Wu, Yao-Hung Hubert Tsai +3
Measuring divergence between two distributions is essential in machine learning and statistics and has various applications including binary classification, change point detection,…