5 papers
BenCSSmark: Making the Social Sciences Count in LLM Research
Arnault Chatelain, Ãtienne Ollion, Qianwen Guan +7
This position paper argues that the under-representation of social science tasks in contemporary LLM benchmarks limits advances in both LLM evaluation and social scientific inquiry…
Pantagruel: Unified Self-Supervised Encoders for French Text and Speech
Phuong-Hang Le, Valentin Pelloin, Arnault Chatelain +27
We release Pantagruel models, a new family of self-supervised encoder models for French text and speech. Instead of predicting modality-tailored targets such as textual tokens or s…
Modeling the human lexicon under temperature variations: linguistic factors, diversity and typicality in LLM word associations
Maria Andueza Rodriguez, Marie Candito, Richard Huyghe
Large language models (LLMs) achieve impressive results in terms of fluency in text generation, yet the nature of their linguistic knowledge - in particular the human-likeness of t…
In the LLM era, Word Sense Induction remains unsolved
Anna Mosolova, Marie Candito, Carlos Ramisch
In the absence of sense-annotated data, word sense induction (WSI) is a compelling alternative to word sense disambiguation, particularly in low-resource or domain-specific setting…
The Self-Contained Negation Test Set
David Kletz, Pascal Amsili, Marie Candito
Several methodologies have recently been proposed to evaluate the ability of Pretrained Language Models (PLMs) to interpret negation. In this article, we build on Gubelmann and Han…