33 citations · 33 across the 1 of their papers we have counts for
1 paper
Javier de la Rosa, Eduardo G. Ponferrada, Paulo Villegas +3
The pre-training of large language models usually requires massive amounts of resources, both in terms of computation and data. Frequently used web sources such as Common Crawl mig…