1 paper
Yajiao Liu, Congliang Chen, Junchi Yang +1
Training large language models with data collected from various domains can improve their performance on downstream tasks. However, given a fixed training budget, the sampling prop…