2 papers
cs.LG2026
When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks
Berk Tinaz, Changzhi Xie, Mahdi Soltanolkotabi
In this paper, we study the gradient descent dynamics for jointly training both layers of a one-hidden-layer ReLU network to fit a linear target function. Concretely, we consider a…
cs.LG2023
Implicit Balancing and Regularization: Generalization and Convergence Guarantees for Overparameterized Asymmetric Matrix Sensing
Mahdi Soltanolkotabi, Dominik Stöger, Changzhi Xie
Recently, there has been significant progress in understanding the convergence and generalization properties of gradient-based methods for training overparameterized learning model…