activity
20212026
most citedTournesol: A quest for a large, secure and trustworthy database of reliable human judgments

5 citations · 5 across the 5 of their papers we have counts for

collaborators

6 papers

cs.AI2026

LLMs Prompted for Legal Context Object More: Overrefusal from Small On-Premises LLMs in Criminal Legal Context

Anastasiia Kucherenko, François Brouchoud, Dimitri Percia David +1

While the validity of LLMs' use in the legal context remains subject to ethical and legal debate, legal professionals are already experimenting with personal LLMs, if only for tran…

cs.CL2025

Getting Your Indices in a Row: Full-Text Search for LLM Training Data for Real World

Ines Altemir Marinas, Anastasiia Kucherenko, Alexander Sternfeld +1

The performance of Large Language Models (LLMs) is determined by their training data. Despite the proliferation of open-weight LLMs, access to LLM training data has remained limite…

cs.CL2025

Apertus: Democratizing Open and Compliant LLMs for Global Language Environments

Project Apertus, Alejandro Hernández-Cano, Alexander Hägele +100

We present Apertus, a fully open suite of large language models (LLMs) designed to address two systemic shortcomings in today's open model ecosystem: data compliance and multilingu…

cs.CL2025

Going over Fine Web with a Fine-Tooth Comb: Technical Report of Indexing Fine Web for Problematic Content Search and Retrieval

Inés Altemir Marinas, Anastasiia Kucherenko, Andrei Kucharavy

Large language models (LLMs) rely heavily on web-scale datasets like Common Crawl, which provides over 80\% of training data for some modern models. However, the indiscriminate nat…

cs.CL2025

Low-Perplexity LLM-Generated Sequences and Where To Find Them

Arthur Wuhrmann, Anastasiia Kucherenko, Andrei Kucharavy

As Large Language Models (LLMs) become increasingly widespread, understanding how specific training data shapes their outputs is crucial for transparency, accountability, privacy,…

cs.HC2021★ 5 cited

Tournesol: A quest for a large, secure and trustworthy database of reliable human judgments

Lê-Nguyên Hoang, Louis Faucon, Aidan Jungo +13

Today's large-scale algorithms have become immensely influential, as they recommend and moderate the content that billions of humans are exposed to on a daily basis. They are the d…