8 papers
Multi-Hop Knowledge Composition is Bound by Pretraining Exposure
Yannis Karmim, Luis Marti, Djamé Seddah +1
Large Language Models fail at implicit multi-hop reasoning: a model answers "When was born?" and "Who is 's closest friend?" correctly but fails on "When was 's closest f…
MultiLinguahah : A New Unsupervised Multilingual Acoustic Laughter Segmentation Method
Sofia Callejas, Nahuel Gomez, Catherine Pelachaud +2
Laughter is a social non-vocalization that is universal across cultures and languages, and is crucial for human communication, including social bonding and communication signaling.…
A Simple Method to Enhance Pre-trained Language Models with Speech Tokens for Classification
Nicolas Calbucura, Jose Guillen, Valentin Barriere
This paper presents a simple method that allows to easily enhance textual pre-trained large language models with speech information, when fine-tuned for a specific classification t…
Leveraging Wikidata for Geographically Informed Sociocultural Bias Dataset Creation: Application to Latin America
Yannis Karmim, Renato Pino, Hernan Contreras +6
Large Language Models (LLMs) exhibit inequalities with respect to various cultural contexts. Most prominent open-weights models are trained on Global North data and show prejudicia…
Constructing a Real-World Benchmark for Early Wildfire Detection with the New PYRONEAR-2025 Dataset
Mateo Lostanlen, Nicolas Isla, Jose Guillen +4
Early wildfire detection (EWD) is of the utmost importance to enable rapid response efforts, and thus minimize the negative impacts of wildfire spreads. To this end, we present PYR…
StandUp4AI: A New Multilingual Dataset for Humor Detection in Stand-up Comedy Videos
Valentin Barriere, Nahuel Gomez, Leo Hemamou +2
Aiming towards improving current computational models of humor detection, we propose a new multimodal dataset of stand-up comedies, in seven languages: English, French, Spanish, It…