activity
20182025
most citedMLS: A Large-Scale Multilingual Dataset for Speech Research

354 citations · 364 across the 8 of their papers we have counts for

collaborators
Showing cs.CLShow all

7 papers · 1 filter

cs.CL2025

Omnilingual ASR: Open-Source Multilingual Speech Recognition for 1600+ Languages

Omnilingual ASR team, Gil Keren, Artyom Kozhevnikov +30

Automatic speech recognition (ASR) has advanced in high-resource languages, but most of the world's 7,000+ languages remain unsupported, leaving thousands of long-tail languages be…

cs.CL2025

Effects of Speaker Count, Duration, and Accent Diversity on Zero-Shot Accent Robustness in Low-Resource ASR

Zheng-Xin Yong, Vineel Pratap, Michael Auli +1

To build an automatic speech recognition (ASR) system that can serve everyone in the world, the ASR needs to be robust to a wide range of accents including unseen accents. We syste…

cs.CL2024

Improving Multilingual ASR in the Wild Using Simple N-best Re-ranking

Brian Yan, Vineel Pratap, Shinji Watanabe +1

Multilingual Automatic Speech Recognition (ASR) models are typically evaluated in a setting where the ground-truth language of the speech utterance is known, however, this is often…

cs.CL2021

Parallel Composition of Weighted Finite-State Transducers

Shubho Sengupta, Vineel Pratap, Awni Hannun

Finite-state transducers (FSTs) are frequently used in speech recognition. Transducer composition is an essential operation for combining different sources of information at differ…

cs.CL2020

Scaling Up Online Speech Recognition Using ConvNets

Vineel Pratap, Qiantong Xu, Jacob Kahn +6

We design an online end-to-end speech recognition system based on Time-Depth Separable (TDS) convolutions and Connectionist Temporal Classification (CTC). We improve the core TDS a…

cs.CL2019

End-to-end ASR: from Supervised to Semi-Supervised Learning with Modern Architectures

Gabriel Synnaeve, Qiantong Xu, Jacob Kahn +6

We study pseudo-labeling for the semi-supervised training of ResNet, Time-Depth Separable ConvNets, and Transformers for speech recognition, with either CTC or Seq2Seq loss functio…