activity
20152022
most citedNeural Sequence-to-Sequence Speech Synthesis Using a Hidden Semi-Markov Model Based Structured Attention Mechanism

4 citations · 9 across the 6 of their papers we have counts for

collaborators

16 papers

eess.AS20221 cited

Embedding a Differentiable Mel-cepstral Synthesis Filter to a Neural Speech Synthesis System

Takenori Yoshimura, Shinji Takaki, Kazuhiro Nakamura +5

This paper integrates a classic mel-cepstral synthesis filter into a modern neural speech synthesis system towards end-to-end controllable speech synthesis. Since the mel-cepstral…

eess.AS20214 cited

Neural Sequence-to-Sequence Speech Synthesis Using a Hidden Semi-Markov Model Based Structured Attention Mechanism

Yoshihiko Nankaku, Kenta Sumiya, Takenori Yoshimura +4

This paper proposes a novel Sequence-to-Sequence (Seq2Seq) model integrating the structure of Hidden Semi-Markov Models (HSMMs) into its attention mechanism. In speech synthesis, i…

eess.AS20211 cited

PeriodNet: A non-autoregressive waveform generation model with a structure separating periodic and aperiodic components

Yukiya Hono, Shinji Takaki, Kei Hashimoto +3

We propose PeriodNet, a non-autoregressive (non-AR) waveform generation model with a new model structure for modeling periodic and aperiodic components in speech waveforms. The non…

cs.SD2019

Transformation of low-quality device-recorded speech to high-quality speech using improved SEGAN model

Seyyed Saeed Sarfjoo, Xin Wang, Gustav Eje Henter +3

Nowadays vast amounts of speech data are recorded from low-quality recorder devices such as smartphones, tablets, laptops, and medium-quality microphones. The objective of this res…

eess.AS2019

Modeling of Rakugo Speech and Its Limitations: Toward Speech Synthesis That Entertains Audiences

Shuhei Kato, Yusuke Yasuda, Xin Wang +3

We have been investigating rakugo speech synthesis as a challenging example of speech synthesis that entertains audiences. Rakugo is a traditional Japanese form of verbal entertain…

eess.AS2019

Fast and High-Quality Singing Voice Synthesis System based on Convolutional Neural Networks

Kazuhiro Nakamura, Shinji Takaki, Kei Hashimoto +3

The present paper describes singing voice synthesis based on convolutional neural networks (CNNs). Singing voice synthesis systems based on deep neural networks (DNNs) are currentl…