1 paper
Cheng-Wei Lin, Wan-Hsuan Hsieh, Kai-Xin Guan +6
The quality and size of a pretraining dataset significantly influence the performance of large language models (LLMs). While there have been numerous efforts in the curation of suc…