11 papers
Average Gradient Outer Product in kernel regression provably recovers the central subspace for multi-index models
Libin Zhu, Damek Davis, Dmitriy Drusvyatskiy +1
We study a prototypical situation when a learned predictor can discover useful low-dimensional structure in data, while using fewer samples than are needed for accurate prediction.…
Forbidden subgraphs in divisor graphs and an ErdÅs divisibility problem
Damek Davis
ErdÅs asked for the largest size of a subset of with no element dividing two others. We show that for an effectively computable constant…
A short proof of near-linear convergence of adaptive gradient descent under fourth-order growth and convexity
Damek Davis, Dmitriy Drusvyatskiy
Davis, Drusvyatskiy, and Jiang showed that gradient descent with an adaptive stepsize converges locally at a nearly-linear rate for smooth functions that grow at least quartically…
The sharp one-dimensional convex sub-Gaussian comparison constant
Damek Davis, Sam Power
Let be an integrable real random variable with mean zero and two-sided sub-Gaussian tail for all . We determine the smallest consta…
Counting partial Hadamard matrices in the cubic regime
Damek Davis
We give a precise asymptotic formula for the number of partial Hadamard matrices in the regimes and for sufficiently large fixed . Th…
When do spectral gradient updates help in deep learning?
Damek Davis, Dmitriy Drusvyatskiy
Spectral gradient methods, such as the recently popularized Muon optimizer, are a promising alternative to standard Euclidean gradient descent for training deep neural networks and…