17 citations · 19 across the 11 of their papers we have counts for
1 paper · 2 filters
Sohir Maskey, Philipp Scholl, Jonas Knupp +2
Language-model checkpoints are commonly selected by pretraining loss or benchmark scores, assuming that the highest-scoring checkpoint will remain the best starting point for subse…