activity
20242026
collaborators

6 papers

cs.CL2026

Post-training makes large language models less human-like

Marcel Binz, Elif Akata, Abdullah Almaatouq +76

Large language models (LLMs) are increasingly used as surrogates for human participants, but it remains unclear which models best capture human behavior and why. To address this, w…

cs.CY2026

StereoTales: A Multilingual Framework for Open-Ended Stereotype Discovery in LLMs

Pierre Le Jeune, Étienne Duchesne, Weixuan Xiao +4

Multilingual studies of social bias in open-ended LLM generation remain limited: most existing benchmarks are English-centric, template-based, or restricted to recognizing pre-spec…

cs.CL2025

PhyloLM : Inferring the Phylogeny of Large Language Models and Predicting their Performances in Benchmarks

Nicolas Yax, Pierre-Yves Oudeyer, Stefano Palminteri

This paper introduces PhyloLM, a method adapting phylogenetic algorithms to Large Language Models (LLMs) to explore whether and how they relate to each other and to predict their p…

cs.CL2025

LogProber: Disentangling confidence from contamination in LLM responses

Nicolas Yax, Pierre-Yves Oudeyer, Stefano Palminteri

In machine learning, contamination refers to situations where testing data leak into the training set. The issue is particularly relevant for the evaluation of the performance of L…

cs.HC2024

The Moral Turing Test: Evaluating Human-LLM Alignment in Moral Decision-Making

Basile Garcia, Crystal Qian, Stefano Palminteri

As large language models (LLMs) become increasingly integrated into society, their alignment with human morals is crucial. To better understand this alignment, we created a large c…

cs.CL2024

Large Language Models are Biased Reinforcement Learners

William M. Hayes, Nicolas Yax, Stefano Palminteri

In-context learning enables large language models (LLMs) to perform a variety of tasks, including learning to make reward-maximizing choices in simple bandit tasks. Given their pot…