3 citations · 8 across the 4 of their papers we have counts for
1 paper · 1 filter
Greg Yang, Etai Littwin
Going beyond stochastic gradient descent (SGD), what new phenomena emerge in wide neural networks trained by adaptive optimizers like Adam? Here we show: The same dichotomy between…