4 papers
Post-training makes large language models less human-like
Marcel Binz, Elif Akata, Abdullah Almaatouq +76
Large language models (LLMs) are increasingly used as surrogates for human participants, but it remains unclear which models best capture human behavior and why. To address this, w…
StereoTales: A Multilingual Framework for Open-Ended Stereotype Discovery in LLMs
Pierre Le Jeune, Ãtienne Duchesne, Weixuan Xiao +4
Multilingual studies of social bias in open-ended LLM generation remain limited: most existing benchmarks are English-centric, template-based, or restricted to recognizing pre-spec…
PhyloLM : Inferring the Phylogeny of Large Language Models and Predicting their Performances in Benchmarks
Nicolas Yax, Pierre-Yves Oudeyer, Stefano Palminteri
This paper introduces PhyloLM, a method adapting phylogenetic algorithms to Large Language Models (LLMs) to explore whether and how they relate to each other and to predict their p…
LogProber: Disentangling confidence from contamination in LLM responses
Nicolas Yax, Pierre-Yves Oudeyer, Stefano Palminteri
In machine learning, contamination refers to situations where testing data leak into the training set. The issue is particularly relevant for the evaluation of the performance of L…