5 papers
On The Landscape of Spoken Language Models: A Comprehensive Survey
Siddhant Arora, Kai-Wei Chang, Chung-Ming Chien +7
The field of spoken language processing is undergoing a shift from training custom-built, task-specific models toward using and optimizing spoken language models (SLMs) which act a…
DiscoPhon: Benchmarking the Unsupervised Discovery of Phoneme Inventories With Discrete Speech Units
Maxime Poli, Manel Khentout, Angelo Ortiz Tandazo +3
We introduce DiscoPhon, a multilingual benchmark for evaluating unsupervised phoneme discovery from discrete speech units. DiscoPhon covers 6 dev and 6 test languages, chosen to sp…
Linguistically Informed Evaluation of Multilingual ASR for African Languages
Fei-Yueh Chen, Lateef Adeleke, C. M. Downey
Word Error Rate (WER) mischaracterizes ASR models' performance for African languages by combining phonological, tone, and other linguistic errors into a single lexical error. By co…
SpidR: Learning Fast and Stable Linguistic Units for Spoken Language Models Without Supervision
Maxime Poli, Mahi Luthra, Youssef Benchekroun +8
The parallel advances in language modeling and speech representation learning have raised the prospect of learning language directly from speech without textual intermediates. This…
LongTail-Swap: benchmarking language models' abilities on rare words
Robin Algayres, Charles-Ãric Saint-James, Mahi Luthra +6
Children learn to speak with a low amount of data and can be taught new words on a few-shot basis, making them particularly data-efficient learners. The BabyLM challenge aims at ex…