1 paper
Yasaman Bahri, Ethan Dyer, Jared Kaplan +2
The population loss of trained deep neural networks often follows precise power-law scaling relations with either the size of the training dataset or the number of parameters in th…