collaborators

6 papers

cs.CL2026

Common Corpus: The Largest Collection of Ethical Data for LLM Pre-Training

Pierre-Carl Langlais, Pavel Chizhov, Catherine Arnett +7

Large Language Models (LLMs) are pre-trained on large amounts of data from different sources and domains. Such datasets often contain trillions of tokens, including large portions…

cs.CL2026

Model in Distress: Sentiment Analysis on French Synthetic Social Media

Pierre-Carl Langlais, Pavel Chizhov, Yannick Detrois +3

Automated analysis of customer feedback on social media is hindered by three challenges: the high cost of annotated training data, the scarcity of evaluation sets, especially in mu…

cs.IR2025

From Show Programmes to Data: Designing a Workflow to Make Performing Arts Ephemera Accessible Through Language Models

Clarisse Bardiot, Pierre-Carl Langlais, Bernard Jacquemin +5

Many heritage institutions hold extensive collections of theatre programmes, which remain largely underused due to their complex layouts and lack of structured metadata. In this pa…

cs.CL2025

Even Small Reasoners Should Quote Their Sources: Introducing the Pleias-RAG Model Family

Pierre-Carl Langlais, Pavel Chizhov, Mattia Nee +6

We introduce a new generation of small reasoning models for RAG, search, and source summarization. Pleias-RAG-350m and Pleias-RAG-1B are mid-trained on a large synthetic dataset em…

cs.CL2025

What the HellaSwag? On the Validity of Common-Sense Reasoning Benchmarks

Pavel Chizhov, Mattia Nee, Pierre-Carl Langlais +1

Common-sense reasoning is a key language model capability because it encapsulates not just specific factual knowledge but rather general language and world understanding. Measuring…

cs.CY2025

Towards Best Practices for Open Datasets for LLM Training

Stefan Baack, Stella Biderman, Kasia Odrozek +36

Many AI companies are training their large language models (LLMs) on data without the permission of the copyright owners. The permissibility of doing so varies by jurisdiction: in…