7 papers
Which Data Matter? Embedding-Based Data Selection for Speech Recognition
Zakaria Aldeneh, Skyler Seto, Maureen de Seyssel +8
Modern ASR systems are typically trained on large-scale pseudo-labeled, in-the-wild data spanning multiple domains. While such heterogeneous data benefit generalist models designed…
Closing the Gap Between Text and Speech Understanding in LLMs
Santiago Cuervo, Skyler Seto, Maureen de Seyssel +5
Large Language Models (LLMs) can be adapted to extend their text capabilities to speech inputs. However, these speech-adapted LLMs consistently underperform their text-based counte…
Discriminating Form and Meaning in Multilingual Models with Minimal-Pair ABX Tasks
Maureen de Seyssel, Jie Chi, Skyler Seto +3
We introduce a set of training-free ABX-style discrimination tasks to evaluate how multilingual language models represent language identity (form) and semantic content (meaning). I…
Leveraging Audio-Visual Data to Reduce the Multilingual Gap in Self-Supervised Speech Models
MarÃa Andrea Cruz Blandón, Zakaria Aldeneh, Jie Chi +1
Self-supervised learning (SSL) has made significant advances in speech representation learning. Models like wav2vec 2.0 and HuBERT have achieved state-of-the-art results in tasks s…
Toward Machine Interpreting: Lessons from Human Interpreting Studies
Matthias Sperber, Maureen de Seyssel, Jiajun Bao +1
Current speech translation systems, while having achieved impressive accuracies, are rather static in their behavior and do not adapt to real-world situations in ways human interpr…
Assessing the Role of Data Quality in Training Bilingual Language Models
Skyler Seto, Maartje ter Hoeve, Maureen de Seyssel +1
Bilingual and multilingual language models offer a promising path toward scaling NLP systems across diverse languages and users. However, their performance often varies wildly betw…