25 citations · 48 across the 3 of their papers we have counts for
1 paper · 1 filter
Mayee F. Chen, Nicholas Roberts, Kush Bhatia +4
The quality of training data impacts the performance of pre-trained large language models (LMs). Given a fixed budget of tokens, we study how to best select data that leads to good…