activity
20212024
most citedScaling Speech Technology to 1,000+ Languages

116 citations · 118 across the 5 of their papers we have counts for

collaborators

5 papers

cs.SD2024

MusicFlow: Cascaded Flow Matching for Text Guided Music Generation

K R Prajwal, Bowen Shi, Matthew Lee +8

We introduce MusicFlow, a cascaded text-to-music generation model based on flow matching. Based on self-supervised representations to bridge between text descriptions and music aud…

eess.AS2024

Learning Fine-Grained Controllability on Speech Generation via Efficient Fine-Tuning

Chung-Ming Chien, Andros Tjandra, Apoorv Vyas +3

As the scale of generative models continues to grow, efficient reuse and adaptation of pre-trained models have become crucial considerations. In this work, we propose Voicebox Adap…

cs.CL2023116 cited

Scaling Speech Technology to 1,000+ Languages

Vineel Pratap, Andros Tjandra, Bowen Shi +13

Expanding the language coverage of speech technology has the potential to improve access to information for many more people. However, current speech technology is restricted to ab…

cs.CL20232 cited

SpeeChain: A Speech Toolkit for Large-Scale Machine Speech Chain

Heli Qi, Sashi Novitasari, Andros Tjandra +2

This paper introduces SpeeChain, an open-source Pytorch-based toolkit designed to develop the machine speech chain for large-scale use. This first release focuses on the TTS-to-ASR…

cs.CL2021

XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Arun Babu, Changhan Wang, Andros Tjandra +10

This paper presents XLS-R, a large-scale model for cross-lingual speech representation learning based on wav2vec 2.0. We train models with up to 2B parameters on nearly half a mill…