6 papers
Common Corpus: The Largest Collection of Ethical Data for LLM Pre-Training
Pierre-Carl Langlais, Pavel Chizhov, Catherine Arnett +7
Large Language Models (LLMs) are pre-trained on large amounts of data from different sources and domains. Such datasets often contain trillions of tokens, including large portions…
Model in Distress: Sentiment Analysis on French Synthetic Social Media
Pierre-Carl Langlais, Pavel Chizhov, Yannick Detrois +3
Automated analysis of customer feedback on social media is hindered by three challenges: the high cost of annotated training data, the scarcity of evaluation sets, especially in mu…
From Show Programmes to Data: Designing a Workflow to Make Performing Arts Ephemera Accessible Through Language Models
Clarisse Bardiot, Pierre-Carl Langlais, Bernard Jacquemin +5
Many heritage institutions hold extensive collections of theatre programmes, which remain largely underused due to their complex layouts and lack of structured metadata. In this pa…
Even Small Reasoners Should Quote Their Sources: Introducing the Pleias-RAG Model Family
Pierre-Carl Langlais, Pavel Chizhov, Mattia Nee +6
We introduce a new generation of small reasoning models for RAG, search, and source summarization. Pleias-RAG-350m and Pleias-RAG-1B are mid-trained on a large synthetic dataset em…
What the HellaSwag? On the Validity of Common-Sense Reasoning Benchmarks
Pavel Chizhov, Mattia Nee, Pierre-Carl Langlais +1
Common-sense reasoning is a key language model capability because it encapsulates not just specific factual knowledge but rather general language and world understanding. Measuring…
Towards Best Practices for Open Datasets for LLM Training
Stefan Baack, Stella Biderman, Kasia Odrozek +36
Many AI companies are training their large language models (LLMs) on data without the permission of the copyright owners. The permissibility of doing so varies by jurisdiction: in…