activity
20192022
most citedESPnet2-TTS: Extending the Edge of TTS Research

28 citations · 45 across the 4 of their papers we have counts for

collaborators

5 papers

eess.AS20221 cited

Embedding a Differentiable Mel-cepstral Synthesis Filter to a Neural Speech Synthesis System

Takenori Yoshimura, Shinji Takaki, Kazuhiro Nakamura +5

This paper integrates a classic mel-cepstral synthesis filter into a modern neural speech synthesis system towards end-to-end controllable speech synthesis. Since the mel-cepstral…

cs.CL202128 cited

ESPnet2-TTS: Extending the Edge of TTS Research

Tomoki Hayashi, Ryuichi Yamamoto, Takenori Yoshimura +7

This paper describes ESPnet2-TTS, an end-to-end text-to-speech (E2E-TTS) toolkit. ESPnet2-TTS extends our earlier version, ESPnet-TTS, by adding many new features, including: on-th…

eess.AS20214 cited

Neural Sequence-to-Sequence Speech Synthesis Using a Hidden Semi-Markov Model Based Structured Attention Mechanism

Yoshihiko Nankaku, Kenta Sumiya, Takenori Yoshimura +4

This paper proposes a novel Sequence-to-Sequence (Seq2Seq) model integrating the structure of Hidden Semi-Markov Models (HSMMs) into its attention mechanism. In speech synthesis, i…

eess.AS2020

End-to-End Automatic Speech Recognition Integrated With CTC-Based Voice Activity Detection

Takenori Yoshimura, Tomoki Hayashi, Kazuya Takeda +1

This paper integrates a voice activity detection (VAD) function with end-to-end automatic speech recognition toward an online speech interface and transcribing very long audio reco…

cs.CL201912 cited

ESPnet-TTS: Unified, Reproducible, and Integratable Open Source End-to-End Text-to-Speech Toolkit

Tomoki Hayashi, Ryuichi Yamamoto, Katsuki Inoue +6

This paper introduces a new end-to-end text-to-speech (E2E-TTS) toolkit named ESPnet-TTS, which is an extension of the open-source speech processing toolkit ESPnet. The toolkit sup…