1 paper · 1 filter
Chuning Li, Chris J. Maddison
We introduce a predictive model that estimates the pre-training loss of large models from model size (N), batch size (B) and number of weight updates (K). This is the first loss pr…