1 paper
Feng Chen, Daniel Kunin, Atsushi Yamamura +1
In this work, we reveal a strong implicit bias of stochastic gradient descent (SGD) that drives overly expressive networks to much simpler subnetworks, thereby dramatically reducin…