activity
20182023
most citedHigh-dimensional Asymptotics of Feature Learning: How One Gradient Step Improves the Representation

11 citations · 14 across the 4 of their papers we have counts for

collaborators
Showing stat.MLShow all

6 papers · 1 filter

stat.ML202211 cited

High-dimensional Asymptotics of Feature Learning: How One Gradient Step Improves the Representation

Jimmy Ba, Murat A. Erdogdu, Taiji Suzuki +3

We study the first gradient descent step on the first-layer parameters in a two-layer neural network: $f(\boldsymbol{x}) = \frac{1}{\sqrt{N}}\boldsymbol{a}^\topσ(\…

stat.ML20221 cited

Convex Analysis of the Mean Field Langevin Dynamics

Atsushi Nitanda, Denny Wu, Taiji Suzuki

As an example of the nonlinear Fokker-Planck equation, the mean field Langevin dynamics recently attracts attention due to its connection to (noisy) gradient descent on infinitely…

stat.ML2020

When Does Preconditioning Help or Hurt Generalization?

Shun-ichi Amari, Jimmy Ba, Roger Grosse +5

While second order optimizers such as natural gradient descent (NGD) often speed up optimization, their effect on generalization has been called into question. This work presents a…

stat.ML2020

On the Optimal Weighted Regularization in Overparameterized Linear Regression

Denny Wu, Ji Xu

We consider the linear model with in the overparameterized regime . We estimate $\ma…

stat.ML2019

Stochastic Runge-Kutta Accelerates Langevin Monte Carlo and Beyond

Xuechen Li, Denny Wu, Lester Mackey +1

Sampling with Markov chain Monte Carlo methods often amounts to discretizing some continuous-time dynamics with numerical integration. In this paper, we establish the convergence r…

stat.ML2018

Post Selection Inference with Incomplete Maximum Mean Discrepancy Estimator

Makoto Yamada, Denny Wu, Yao-Hung Hubert Tsai +3

Measuring divergence between two distributions is essential in machine learning and statistics and has various applications including binary classification, change point detection,…