1 paper
Berk Tinaz, Changzhi Xie, Mahdi Soltanolkotabi
In this paper, we study the gradient descent dynamics for jointly training both layers of a one-hidden-layer ReLU network to fit a linear target function. Concretely, we consider a…