1 paper
Lior Belenki, Alekh Agarwal, Tianze Shi +1
We propose a method to optimize language model pre-training data mixtures through efficient approximation of the cross-entropy loss corresponding to each candidate mixture via a Mi…