157 citations · 157 across the 1 of their papers we have counts for
1 paper
Guilherme Penedo, Quentin Malartic, Daniel Hesslow +6
Large language models are commonly trained on a mixture of filtered web data and curated high-quality corpora, such as social media conversations, books, or technical papers. This…