10 papers
Recovery of latent inner products from an anisotropic Gaussian random geometric graph
Cheng Mao, Vidya Muthukumar
We study the problem of recovering latent inner products from a random geometric graph with anisotropic Gaussian latent points. More precisely, for an i.i.d. sample $x_1, \dots, x_…
SGD Provably Prioritizes a Shortcut Spurious Feature in the XOR Model
Tyler LaBonte, Vidya Muthukumar
Neural networks are known to be susceptible to over-reliance on spurious correlations. However, the precise mechanism by which models exploit shortcut features is not fully underst…
How Does the ReLU Activation Affect the Implicit Bias of Gradient Descent on High-dimensional Neural Network Regression?
Kuo-Wei Lai, Guanghui Wang, Molei Tao +1
Overparameterized ML models, including neural networks, typically induce underdetermined training objectives with multiple global minima. The implicit bias refers to the limiting g…
On the Unreasonable Effectiveness of Last-layer Retraining
John C. Hill, Tyler LaBonte, Xinchen Zhang +1
Last-layer retraining (LLR) methods -- wherein the last layer of a neural network is reinitialized and retrained on a held-out set following ERM training -- have garnered interest…
Last-iterate Convergence for Symmetric, General-sum, Games Under The Exponential Weights Dynamic
Guanghui Wang, Krishna Acharya, Lokranjan Lakshmikanthan +2
We conduct a comprehensive analysis of the discrete-time exponential-weights dynamic with a constant step size on all general-sum and symmetric normal-form games, i.e.…
A general technique for approximating high-dimensional empirical kernel matrices
Chiraag Kaushik, Justin Romberg, Vidya Muthukumar
We present simple, user-friendly bounds for the expected operator norm of a random kernel matrix under general conditions on the kernel function . Our approach uses…