1 citations · 1 across the 1 of their papers we have counts for
3 papers
Wikidata Search Traces: A Dataset for Training Knowledge Graph Search Agents
Mohamed Chenene, Carlos Rosas-Hinostroza, Pierre-Carl Langlais +2
Wikidata is one of the largest open knowledge bases, yet answering a complex question over it still requires a SPARQL query that names the right entities and properties and chains…
It's All Training: A Fully Synthetic Single-Stage Recipe for LLMs
Pierre-Carl Langlais, Pieter Delobelle, Yannick Detrois +7
Current pre-training datasets are derived from web crawls, with all their issues, and were not designed to support mid- and post-training pipelines--for instance, they contain litt…
Common Corpus: The Largest Collection of Ethical Data for LLM Pre-Training
Pierre-Carl Langlais, Pavel Chizhov, Catherine Arnett +8
Large Language Models (LLMs) are pre-trained on large amounts of data from different sources and domains. Such datasets often contain trillions of tokens, including large portions…