1 paper · 1 filter
Yiming Zhang, Javier Rando, Ivan Evtimov +5
Large language models are pre-trained on uncurated text datasets consisting of trillions of tokens scraped from the Web. Prior work has shown that: (1) web-scraped pre-training dat…