4 citations · 9 across the 5 of their papers we have counts for
1 paper · 1 filter
Ishaan Watts, Catherine Li, Sachin Goyal +2
Pretraining optimizers are tuned to produce the strongest possible base model, on the assumption that a stronger starting point yields a stronger model after subsequent changes lik…