most citedVoice Transformer Network: Sequence-to-Sequence Voice Conversion Using Transformer with Text-to-Speech Pretraining

38 citations · 98 across the 8 of their papers we have counts for

collaborators

12 papers

eess.AS20207 cited

The Sequence-to-Sequence Baseline for the Voice Conversion Challenge 2020: Cascading ASR and TTS

Wen-Chin Huang, Tomoki Hayashi, Shinji Watanabe +1

This paper presents the sequence-to-sequence (seq2seq) baseline system for the voice conversion challenge (VCC) 2020. We consider a naive approach for voice conversion (VC), which…

eess.AS20203 cited

Pretraining Techniques for Sequence-to-Sequence Voice Conversion

Wen-Chin Huang, Tomoki Hayashi, Yi-Chiao Wu +2

Sequence-to-sequence (seq2seq) voice conversion (VC) models are attractive owing to their ability to convert prosody. Nonetheless, without sufficient data, seq2seq VC models can su…

cs.CL202021 cited

DiscreTalk: Text-to-Speech as a Machine Translation Problem

Tomoki Hayashi, Shinji Watanabe

This paper proposes a new end-to-end text-to-speech (E2E-TTS) model based on neural machine translation (NMT). The proposed model consists of two components; a non-autoregressive v…

eess.AS20202 cited

Quasi-Periodic Parallel WaveGAN Vocoder: A Non-autoregressive Pitch-dependent Dilated Convolution Model for Parametric Speech Generation

Yi-Chiao Wu, Tomoki Hayashi, Takuma Okamoto +2

In this paper, we propose a parallel WaveGAN (PWG)-like neural vocoder with a quasi-periodic (QP) architecture to improve the pitch controllability of PWG. PWG is a compact non-aut…

cs.CL2020

ESPnet-ST: All-in-One Speech Translation Toolkit

Hirofumi Inaguma, Shun Kiyono, Kevin Duh +4

We present ESPnet-ST, which is designed for the quick development of speech-to-speech translation systems in a single framework. ESPnet-ST is a new project inside end-to-end speech…

eess.AS2020

End-to-End Automatic Speech Recognition Integrated With CTC-Based Voice Activity Detection

Takenori Yoshimura, Tomoki Hayashi, Kazuya Takeda +1

This paper integrates a voice activity detection (VAD) function with end-to-end automatic speech recognition toward an online speech interface and transcribing very long audio reco…