1 paper
Joshua Shay Kricheli, Alexander Lawrence Reid, Soumajyoti Sarkar +2
Neural scaling laws approximate a language model's loss as a power-law function of parameter count N and token count D. Following Chinchilla-style compute-optimal training, man…