11 citations · 11 across the 1 of their papers we have counts for
1 paper
Jimmy Ba, Murat A. Erdogdu, Taiji Suzuki +3
We study the first gradient descent step on the first-layer parameters W in a two-layer neural network: $f(\boldsymbol{x}) = \frac{1}{\sqrt{N}}\boldsymbol{a}^\topσ(\…