activity
20212023
most citedScaling Speech Technology to 1,000+ Languages

116 citations · 259 across the 8 of their papers we have counts for

collaborators

8 papers

cs.CL2023116 cited

Scaling Speech Technology to 1,000+ Languages

Vineel Pratap, Andros Tjandra, Bowen Shi +13

Expanding the language coverage of speech technology has the potential to improve access to information for many more people. However, current speech technology is restricted to ab…

cs.CL20232 cited

Cocktail HuBERT: Generalized Self-Supervised Pre-training for Mixture and Single-Source Speech

Maryam Fazel-Zarandi, Wei-Ning Hsu

Self-supervised learning leverages unlabeled data effectively, improving label efficiency and generalization to domains without labeled data. While recent work has studied generali…

cs.CL20234 cited

MuAViC: A Multilingual Audio-Visual Corpus for Robust Speech Recognition and Robust Speech-to-Text Translation

Mohamed Anwar, Bowen Shi, Vedanuj Goswami +3

We introduce MuAViC, a multilingual audio-visual corpus for robust speech recognition and robust speech-to-text translation providing 1200 hours of audio-visual speech in 9 languag…

cs.CL20237 cited

Scaling Laws for Generative Mixed-Modal Language Models

Armen Aghajanyan, Lili Yu, Alexis Conneau +7

Generative language models define distributions over sequences of tokens that can represent essentially any combination of data modalities (e.g., any permutation of image tokens fr…

cs.CL2022

STOP: A dataset for Spoken Task Oriented Semantic Parsing

Paden Tomasello, Akshat Shrivastava, Daniel Lazar +12

End-to-end spoken language understanding (SLU) predicts intent directly from audio using a single model. It promises to improve the performance of assistant systems by leveraging a…

cs.CL202216 cited

u-HuBERT: Unified Mixed-Modal Speech Pretraining And Zero-Shot Transfer to Unlabeled Modality

Wei-Ning Hsu, Bowen Shi

While audio-visual speech models can yield superior performance and robustness compared to audio-only models, their development and adoption are hindered by the lack of labeled and…