1 citations · 1 across the 1 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026★ 1 cited
Common Corpus: The Largest Collection of Ethical Data for LLM Pre-Training
Pierre-Carl Langlais, Pavel Chizhov, Catherine Arnett +7
Large Language Models (LLMs) are pre-trained on large amounts of data from different sources and domains. Such datasets often contain trillions of tokens, including large portions…
cs.CL2024
Toxicity of the Commons: Curating Open-Source Pre-Training Data
Catherine Arnett, Eliot Jones, Ivan P. Yamshchikov +1
Open-source large language models are becoming increasingly available and popular among researchers and practitioners. While significant progress has been made on open-weight model…