2 papers
cs.CL2026
It's All Training: A Fully Synthetic Single-Stage Recipe for LLMs
Pierre-Carl Langlais, Pieter Delobelle, Yannick Detrois +7
Current pre-training datasets are derived from web crawls, with all their issues, and were not designed to support mid- and post-training pipelines--for instance, they contain litt…
cs.CL2025
When Scale Meets Diversity: Evaluating Language Models on Fine-Grained Multilingual Claim Verification
Hanna Shcharbakova, Tatiana Anikina, Natalia Skachkova +1
The rapid spread of multilingual misinformation requires robust automated fact verification systems capable of handling fine-grained veracity assessments across diverse languages.…