1 paper · 1 filter
John Sous, Michael Winer
We introduce a model for neural scaling laws under sparse activations. In the model, test loss is often dominated by rare coordinates that are never observed in the training input.…