54 citations · 161 across the 6 of their papers we have counts for
1 paper · 1 filter
Greg Yang, Etai Littwin
Going beyond stochastic gradient descent (SGD), what new phenomena emerge in wide neural networks trained by adaptive optimizers like Adam? Here we show: The same dichotomy between…