activity
20132022
most citedLingvo: a Modular and Scalable Framework for Sequence-to-Sequence Modeling

184 citations · 285 across the 11 of their papers we have counts for

collaborators
Showing cs.CLShow all

7 papers · 1 filter

cs.CL2020

Wave-Tacotron: Spectrogram-free end-to-end text-to-speech synthesis

Ron J. Weiss, RJ Skerry-Ryan, Eric Battenberg +2

We describe a sequence-to-sequence neural network which directly generates speech waveforms from text inputs. The architecture extends the Tacotron model by incorporating a normali…

cs.CL201926 cited

Learning to Speak Fluently in a Foreign Language: Multilingual Speech Synthesis and Cross-Language Voice Cloning

Yu Zhang, Ron J. Weiss, Heiga Zen +6

We present a multispeaker, multilingual text-to-speech (TTS) synthesis model based on Tacotron that is able to produce high quality speech in multiple languages. Moreover, the mode…

cs.CL201922 cited

Direct speech-to-speech translation with a sequence-to-sequence model

Ye Jia, Ron J. Weiss, Fadi Biadsy +4

We present an attention-based sequence-to-sequence neural network which can directly translate speech from one language into speech in another language, without relying on an inter…

cs.CL2018

Leveraging Weakly Supervised Data to Improve End-to-End Speech-to-Text Translation

Ye Jia, Melvin Johnson, Wolfgang Macherey +6

End-to-end Speech Translation (ST) models have many potential advantages when compared to the cascade of Automatic Speech Recognition (ASR) and text Machine Translation (MT) models…

cs.CL2018

Hierarchical Generative Modeling for Controllable Speech Synthesis

Wei-Ning Hsu, Yu Zhang, Ron J. Weiss +9

This paper proposes a neural sequence-to-sequence text-to-speech (TTS) model which can control latent attributes in the generated speech that are rarely annotated in the training d…

cs.CL2018

Transfer Learning from Speaker Verification to Multispeaker Text-To-Speech Synthesis

Ye Jia, Yu Zhang, Ron J. Weiss +8

We describe a neural network-based system for text-to-speech (TTS) synthesis that is able to generate speech audio in the voice of many different speakers, including those unseen d…