activity
20212024
most citedScaling Speech Technology to 1,000+ Languages

116 citations · 261 across the 12 of their papers we have counts for

collaborators
Showing cs.CLShow all

9 papers · 1 filter

cs.CL20232 cited

Toward Joint Language Modeling for Speech Units and Text

Ju-Chieh Chou, Chung-Ming Chien, Wei-Ning Hsu +5

Speech and text are two major forms of human language. The research community has been focusing on mapping speech to text or vice versa for many years. However, in the field of lan…

cs.CL2023

EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Tu Anh Nguyen, Wei-Ning Hsu, Antony D'Avirro +10

Recent work has shown that it is possible to resynthesize high-quality speech based, not on text, but on low bitrate discrete units that have been learned in a self-supervised fash…

cs.CL2023116 cited

Scaling Speech Technology to 1,000+ Languages

Vineel Pratap, Andros Tjandra, Bowen Shi +13

Expanding the language coverage of speech technology has the potential to improve access to information for many more people. However, current speech technology is restricted to ab…

cs.CL20232 cited

Cocktail HuBERT: Generalized Self-Supervised Pre-training for Mixture and Single-Source Speech

Maryam Fazel-Zarandi, Wei-Ning Hsu

Self-supervised learning leverages unlabeled data effectively, improving label efficiency and generalization to domains without labeled data. While recent work has studied generali…

cs.CL20234 cited

MuAViC: A Multilingual Audio-Visual Corpus for Robust Speech Recognition and Robust Speech-to-Text Translation

Mohamed Anwar, Bowen Shi, Vedanuj Goswami +3

We introduce MuAViC, a multilingual audio-visual corpus for robust speech recognition and robust speech-to-text translation providing 1200 hours of audio-visual speech in 9 languag…

cs.CL20237 cited

Scaling Laws for Generative Mixed-Modal Language Models

Armen Aghajanyan, Lili Yu, Alexis Conneau +7

Generative language models define distributions over sequences of tokens that can represent essentially any combination of data modalities (e.g., any permutation of image tokens fr…