activity
20162022
most citedLearning to Speak Fluently in a Foreign Language: Multilingual Speech Synthesis and Cross-Language Voice Cloning

26 citations · 85 across the 14 of their papers we have counts for

collaborators
Showing cs.CLShow all

10 papers · 1 filter

cs.CL2022

Maestro-U: Leveraging joint speech-text representation learning for zero supervised speech ASR

Zhehuai Chen, Ankur Bapna, Andrew Rosenberg +4

Training state-of-the-art Automated Speech Recognition (ASR) models typically requires a substantial amount of transcribed speech. In this work, we demonstrate that a modality-matc…

cs.CL2022

Analysis of Self-Attention Head Diversity for Conformer-based Automatic Speech Recognition

Kartik Audhkhasi, Yinghui Huang, Bhuvana Ramabhadran +1

Attention layers are an integral part of modern end-to-end automatic speech recognition systems, for instance as part of the Transformer or Conformer architecture. Attention is typ…

cs.CL20214 cited

Injecting Text in Self-Supervised Speech Pretraining

Zhehuai Chen, Yu Zhang, Andrew Rosenberg +3

Self-supervised pretraining for Automated Speech Recognition (ASR) has shown varied degrees of success. In this paper, we propose to jointly learn representations during pretrainin…

cs.CL2019

Speech Recognition with Augmented Synthesized Speech

Andrew Rosenberg, Yu Zhang, Bhuvana Ramabhadran +4

Recent success of the Tacotron speech synthesis architecture and its variants in producing natural sounding multi-speaker synthesized speech has raised the exciting possibility of…

cs.CL201926 cited

Learning to Speak Fluently in a Foreign Language: Multilingual Speech Synthesis and Cross-Language Voice Cloning

Yu Zhang, Ron J. Weiss, Heiga Zen +6

We present a multispeaker, multilingual text-to-speech (TTS) synthesis model based on Tacotron that is able to produce high quality speech in multiple languages. Moreover, the mode…

cs.CL2018

Joint Modeling of Accents and Acoustics for Multi-Accent Speech Recognition

Xuesong Yang, Kartik Audhkhasi, Andrew Rosenberg +3

The performance of automatic speech recognition systems degrades with increasing mismatch between the training and testing scenarios. Differences in speaker accents are a significa…