88 citations · 127 across the 36 of their papers we have counts for
6 papers · 1 filter
JSSS: free Japanese speech corpus for summarization and simplification
Shinnosuke Takamichi, Mamoru Komachi, Naoko Tanji +1
In this paper, we construct a new Japanese speech corpus for speech-based summarization and simplification, "JSSS" (pronounced "j-triple-s"). Given the success of reading-style spe…
Multi-speaker Text-to-speech Synthesis Using Deep Gaussian Processes
Kentaro Mitsui, Tomoki Koriyama, Hiroshi Saruwatari
Multi-speaker speech synthesis is a technique for modeling multiple speakers' voices with a single model. Although many approaches using deep neural networks (DNNs) have been propo…
Utterance-level Sequential Modeling For Deep Gaussian Process Based Speech Synthesis Using Simple Recurrent Unit
Tomoki Koriyama, Hiroshi Saruwatari
This paper presents a deep Gaussian process (DGP) model with a recurrent architecture for speech sequence modeling. DGP is a Bayesian deep model that can be trained effectively wit…
DNN-based Speaker Embedding Using Subjective Inter-speaker Similarity for Multi-speaker Modeling in Speech Synthesis
Yuki Saito, Shinnosuke Takamichi, Hiroshi Saruwatari
This paper proposes novel algorithms for speaker embedding using subjective inter-speaker similarity based on deep neural networks (DNNs). Although conventional DNN-based speaker e…
Independent Low-Rank Matrix Analysis Based on Time-Variant Sub-Gaussian Source Model
Shinichi Mogami, Norihiro Takamune, Daichi Kitamura +5
Independent low-rank matrix analysis (ILRMA) is a fast and stable method for blind audio source separation. Conventional ILRMAs assume time-variant (super-)Gaussian source models,…
Independent Deeply Learned Matrix Analysis for Multichannel Audio Source Separation
Shinichi Mogami, Hayato Sumino, Daichi Kitamura +4
In this paper, we address a multichannel audio source separation task and propose a new efficient method called independent deeply learned matrix analysis (IDLMA). IDLMA estimates…