activity
20172024
most citedJSUT corpus: free large-scale Japanese speech corpus for end-to-end speech synthesis

88 citations · 197 across the 76 of their papers we have counts for

collaborators
Showing 2020 · cs.SDShow all

7 papers · 2 filters

cs.SD2020

Incremental Text-to-Speech Synthesis Using Pseudo Lookahead with Large Pretrained Language Model

Takaaki Saeki, Shinnosuke Takamichi, Hiroshi Saruwatari

This letter presents an incremental text-to-speech (TTS) method that performs synthesis in small linguistic units while maintaining the naturalness of output speech. Incremental TT…

cs.SD2020★ 4 cited

Joint-Diagonalizability-Constrained Multichannel Nonnegative Matrix Factorization Based on Multivariate Complex Sub-Gaussian Distribution

Keigo Kamo, Yuki Kubo, Norihiro Takamune +4

In this paper, we address a statistical model extension of multichannel nonnegative matrix factorization (MNMF) for blind source separation, and we propose a new parameter update a…

cs.SD2020

Convergence-guaranteed Independent Positive Semidefinite Tensor Analysis Based on Student's t Distribution

Tatsuki Kondo, Kanta Fukushige, Norihiro Takamune +4

In this paper, we address a blind source separation (BSS) problem and propose a new extended framework of independent positive semidefinite tensor analysis (IPSDTA). IPSDTA is a st…

cs.SD2020

Lifter Training and Sub-band Modeling for Computationally Efficient and High-Quality Voice Conversion Using Spectral Differentials

Takaaki Saeki, Yuki Saito, Shinnosuke Takamichi +1

In this paper, we propose computationally efficient and high-quality methods for statistical voice conversion (VC) with direct waveform modification based on spectral differentials…

cs.SD2020

Regularized Fast Multichannel Nonnegative Matrix Factorization with ILRMA-based Prior Distribution of Joint-Diagonalization Process

Keigo Kamo, Yuki Kubo, Norihiro Takamune +4

In this paper, we address a convolutive blind source separation (BSS) problem and propose a new extended framework of FastMNMF by introducing prior information for joint diagonaliz…

cs.SD2020

Time-Domain Audio Source Separation Based on Wave-U-Net Combined with Discrete Wavelet Transform

Tomohiko Nakamura, Hiroshi Saruwatari

We propose a time-domain audio source separation method using down-sampling (DS) and up-sampling (US) layers based on a discrete wavelet transform (DWT). The proposed method is bas…