activity
20242026
collaborators

8 papers

cs.CL2026

Multi-Hop Knowledge Composition is Bound by Pretraining Exposure

Yannis Karmim, Luis Marti, Djamé Seddah +1

Large Language Models fail at implicit multi-hop reasoning: a model answers "When was born?" and "Who is 's closest friend?" correctly but fails on "When was 's closest f…

cs.CL2026

MultiLinguahah : A New Unsupervised Multilingual Acoustic Laughter Segmentation Method

Sofia Callejas, Nahuel Gomez, Catherine Pelachaud +2

Laughter is a social non-vocalization that is universal across cultures and languages, and is crucial for human communication, including social bonding and communication signaling.…

cs.CL2026

A Simple Method to Enhance Pre-trained Language Models with Speech Tokens for Classification

Nicolas Calbucura, Jose Guillen, Valentin Barriere

This paper presents a simple method that allows to easily enhance textual pre-trained large language models with speech information, when fine-tuned for a specific classification t…

cs.CL2026

Leveraging Wikidata for Geographically Informed Sociocultural Bias Dataset Creation: Application to Latin America

Yannis Karmim, Renato Pino, Hernan Contreras +6

Large Language Models (LLMs) exhibit inequalities with respect to various cultural contexts. Most prominent open-weights models are trained on Global North data and show prejudicia…

cs.CV2025

Constructing a Real-World Benchmark for Early Wildfire Detection with the New PYRONEAR-2025 Dataset

Mateo Lostanlen, Nicolas Isla, Jose Guillen +4

Early wildfire detection (EWD) is of the utmost importance to enable rapid response efforts, and thus minimize the negative impacts of wildfire spreads. To this end, we present PYR…

cs.CL2025

StandUp4AI: A New Multilingual Dataset for Humor Detection in Stand-up Comedy Videos

Valentin Barriere, Nahuel Gomez, Leo Hemamou +2

Aiming towards improving current computational models of humor detection, we propose a new multimodal dataset of stand-up comedies, in seven languages: English, French, Spanish, It…