20 citations · 100 across the 13 of their papers we have counts for
18 papers
Generalization Bounds for Stochastic Gradient Descent via Localized -Covers
Sejun Park, Umut Şimşekli, Murat A. Erdogdu
In this paper, we propose a new covering technique localized for the trajectories of SGD. This localization provides an algorithm-specific complexity measured by the covering numbe…
High-dimensional Asymptotics of Feature Learning: How One Gradient Step Improves the Representation
Jimmy Ba, Murat A. Erdogdu, Taiji Suzuki +3
We study the first gradient descent step on the first-layer parameters in a two-layer neural network: $f(\boldsymbol{x}) = \frac{1}{\sqrt{N}}\boldsymbol{a}^\topσ(\…
Mirror Descent Strikes Again: Optimal Stochastic Convex Optimization under Infinite Noise Variance
Nuri Mert Vural, Lu Yu, Krishnakumar Balasubramanian +2
We study stochastic convex optimization under infinite noise variance. Specifically, when the stochastic gradient is unbiased and has uniformly bounded -th moment, for some…
Towards a Theory of Non-Log-Concave Sampling: First-Order Stationarity Guarantees for Langevin Monte Carlo
Krishnakumar Balasubramanian, Sinho Chewi, Murat A. Erdogdu +2
For the task of sampling from a density on , where is possibly non-convex but -gradient Lipschitz, we prove that averaged Langevin Monte Ca…
Heavy-tailed Sampling via Transformed Unadjusted Langevin Algorithm
Ye He, Krishnakumar Balasubramanian, Murat A. Erdogdu
We analyze the oracle complexity of sampling from polynomially decaying heavy-tailed target densities based on running the Unadjusted Langevin Algorithm on certain transformed vers…
On Empirical Risk Minimization with Dependent and Heavy-Tailed Data
Abhishek Roy, Krishnakumar Balasubramanian, Murat A. Erdogdu
In this work, we establish risk bounds for the Empirical Risk Minimization (ERM) with both dependent and heavy-tailed data-generating processes. We do so by extending the seminal w…