1 paper · 1 filter
Ke Chen, Chugang Yi, Haizhao Yang
We study the implicit bias towards low-rank weight matrices when training neural networks (NN) with Weight Decay (WD). We prove that when a ReLU NN is sufficiently trained with Sto…