Showing cs.CLShow all
3 papers · 1 filter
cs.CL2025
Leveraging Audio-Visual Data to Reduce the Multilingual Gap in Self-Supervised Speech Models
María Andrea Cruz Blandón, Zakaria Aldeneh, Jie Chi +1
Self-supervised learning (SSL) has made significant advances in speech representation learning. Models like wav2vec 2.0 and HuBERT have achieved state-of-the-art results in tasks s…
cs.CL2025
Discriminating Form and Meaning in Multilingual Models with Minimal-Pair ABX Tasks
Maureen de Seyssel, Jie Chi, Skyler Seto +3
We introduce a set of training-free ABX-style discrimination tasks to evaluate how multilingual language models represent language identity (form) and semantic content (meaning). I…
cs.CL2025
The Role of Prosody in Spoken Question Answering
Jie Chi, Maureen de Seyssel, Natalie Schluter
Spoken language understanding research to date has generally carried a heavy text perspective. Most datasets are derived from text, which is then subsequently synthesized into spee…