1 paper
David Meltzer, Min Chen, Junyu Liu
Neural networks trained with gradient descent can undergo non-trivial phase transitions as a function of the learning rate. In \cite{lewkowycz2020large} it was discovered that wide…