collaborators

7 papers

cs.SD2026

Which Data Matter? Embedding-Based Data Selection for Speech Recognition

Zakaria Aldeneh, Skyler Seto, Maureen de Seyssel +8

Modern ASR systems are typically trained on large-scale pseudo-labeled, in-the-wild data spanning multiple domains. While such heterogeneous data benefit generalist models designed…

cs.CL2026

Closing the Gap Between Text and Speech Understanding in LLMs

Santiago Cuervo, Skyler Seto, Maureen de Seyssel +5

Large Language Models (LLMs) can be adapted to extend their text capabilities to speech inputs. However, these speech-adapted LLMs consistently underperform their text-based counte…

cs.CL2025

Discriminating Form and Meaning in Multilingual Models with Minimal-Pair ABX Tasks

Maureen de Seyssel, Jie Chi, Skyler Seto +3

We introduce a set of training-free ABX-style discrimination tasks to evaluate how multilingual language models represent language identity (form) and semantic content (meaning). I…

cs.CL2025

Leveraging Audio-Visual Data to Reduce the Multilingual Gap in Self-Supervised Speech Models

María Andrea Cruz Blandón, Zakaria Aldeneh, Jie Chi +1

Self-supervised learning (SSL) has made significant advances in speech representation learning. Models like wav2vec 2.0 and HuBERT have achieved state-of-the-art results in tasks s…

cs.CL2025

Toward Machine Interpreting: Lessons from Human Interpreting Studies

Matthias Sperber, Maureen de Seyssel, Jiajun Bao +1

Current speech translation systems, while having achieved impressive accuracies, are rather static in their behavior and do not adapt to real-world situations in ways human interpr…

cs.CL2025

Assessing the Role of Data Quality in Training Bilingual Language Models

Skyler Seto, Maartje ter Hoeve, Maureen de Seyssel +1

Bilingual and multilingual language models offer a promising path toward scaling NLP systems across diverse languages and users. However, their performance often varies wildly betw…