1 paper
Yifan Wang, Binbin Liu, Fengze Liu +6
The data mixture used in the pre-training of a language model is a cornerstone of its final performance. However, a static mixing strategy is suboptimal, as the model's learning pr…