activity
20192024
most citedParallel WaveGAN: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram

50 citations · 101 across the 9 of their papers we have counts for

collaborators

10 papers

cs.CL202128 cited

ESPnet2-TTS: Extending the Edge of TTS Research

Tomoki Hayashi, Ryuichi Yamamoto, Takenori Yoshimura +7

This paper describes ESPnet2-TTS, an end-to-end text-to-speech (E2E-TTS) toolkit. ESPnet2-TTS extends our earlier version, ESPnet-TTS, by adding many new features, including: on-th…

eess.AS2021

Phrase break prediction with bidirectional encoder representations in Japanese text-to-speech synthesis

Kosuke Futamata, Byeongseon Park, Ryuichi Yamamoto +1

We propose a novel phrase break prediction method that combines implicit features extracted from a pre-trained large language model, a.k.a BERT, and explicit features extracted fro…

eess.AS20213 cited

Improved parallel WaveGAN vocoder with perceptually weighted spectrogram loss

Eunwoo Song, Ryuichi Yamamoto, Min-Jae Hwang +3

This paper proposes a spectral-domain perceptual weighting technique for Parallel WaveGAN-based text-to-speech (TTS) systems. The recently proposed Parallel WaveGAN vocoder success…

eess.AS20205 cited

TTS-by-TTS: TTS-driven Data Augmentation for Fast and High-Quality Speech Synthesis

Min-Jae Hwang, Ryuichi Yamamoto, Eunwoo Song +1

In this paper, we propose a text-to-speech (TTS)-driven data augmentation method for improving the quality of a non-autoregressive (AR) TTS system. Recently proposed non-AR models,…

eess.AS2020

Parallel waveform synthesis based on generative adversarial networks with voicing-aware conditional discriminators

Ryuichi Yamamoto, Eunwoo Song, Min-Jae Hwang +1

This paper proposes voicing-aware conditional discriminators for Parallel WaveGAN-based waveform synthesis systems. In this framework, we adopt a projection-based conditioning meth…

eess.AS2020

Neural text-to-speech with a modeling-by-generation excitation vocoder

Eunwoo Song, Min-Jae Hwang, Ryuichi Yamamoto +3

This paper proposes a modeling-by-generation (MbG) excitation vocoder for a neural text-to-speech (TTS) system. Recently proposed neural excitation vocoders can realize qualified w…