20 citations · 20 across the 1 of their papers we have counts for
1 paper
Utkarsh Sharma, Jared Kaplan
When data is plentiful, the loss achieved by well-trained neural networks scales as a power-law L∝N−α in the number of network parameters N. This empirical scaling l…