1 paper
Ishaan Watts, Catherine Li, Sachin Goyal +2
Pretraining optimizers are tuned to produce the strongest possible base model, on the assumption that a stronger starting point yields a stronger model after subsequent changes lik…