1 citations · 1 across the 3 of their papers we have counts for
4 papers · 1 filter
SpidR: Learning Fast and Stable Linguistic Units for Spoken Language Models Without Supervision
Maxime Poli, Mahi Luthra, Youssef Benchekroun +8
The parallel advances in language modeling and speech representation learning have raised the prospect of learning language directly from speech without textual intermediates. This…
SpidR-Adapt: A Universal Speech Representation Model for Few-Shot Adaptation
Mahi Luthra, Jiayi Shen, Maxime Poli +14
Human infants, with only a few hundred hours of speech exposure, acquire basic units of new languages, highlighting a striking efficiency gap compared to the data-hungry self-super…
LongTail-Swap: benchmarking language models' abilities on rare words
Robin Algayres, Charles-Éric Saint-James, Mahi Luthra +6
Children learn to speak with a low amount of data and can be taught new words on a few-shot basis, making them particularly data-efficient learners. The BabyLM challenge aims at ex…
WorldSense: A Synthetic Benchmark for Grounded Reasoning in Large Language Models
Youssef Benchekroun, Megi Dervishi, Mark Ibrahim +7
We propose WorldSense, a benchmark designed to assess the extent to which LLMs are consistently able to sustain tacit world models, by testing how they draw simple inferences from…