JSUT corpus: free large-scale Japanese speech corpus for end-to-end speech synthesis
arXiv:1711.00354
Abstract
Thanks to improvements in machine learning techniques including deep learning, a free large-scale speech corpus that can be shared between academic institutions and commercial companies has an important role. However, such a corpus for Japanese speech synthesis does not exist. In this paper, we designed a novel Japanese speech corpus, named the "JSUT corpus," that is aimed at achieving end-to-end speech synthesis. The corpus consists of 10 hours of reading-style speech data and its transcription and covers all of the main pronunciations of daily-use Japanese characters. In this paper, we describe how we designed and analyzed the corpus. The corpus is freely available online.
Submitted to ICASSP2018
Cited by in corpus (17)
- A Survey on Neural Speech Synthesis
- Attack Agnostic Dataset: Towards Generalization and Stabilization of Audio DeepFake Detection
- Adapting Multilingual Speech Representation Model for a New, Underresourced Language through Multilingual Fine-tuning and Continued Pretraining
- ESPnet2-TTS: Extending the Edge of TTS Research
- JVNV: A Corpus of Japanese Emotional Speech with Verbal Content and Nonverbal Expressions
- DiscreTalk: Text-to-Speech as a Machine Translation Problem
- Investigation of Japanese PnG BERT language model in text-to-speech synthesis for pitch accent language
- ESPnet-TTS: Unified, Reproducible, and Integratable Open Source End-to-End Text-to-Speech Toolkit
- Simultaneous Speech-to-Speech Translation System with Neural Incremental ASR, MT, and TTS
- How does a spontaneously speaking conversational agent affect user behavior?
- Unsupervised Learning For Sequence-to-sequence Text-to-speech For Low-resource Languages
- Accent Estimation of Japanese Words from Their Surfaces and Romanizations for Building Large Vocabulary Accent Dictionaries
- RefineGAN: Universally Generating Waveform Better than Ground Truth with Highly Accurate Pitch and Intensity Responses
- JSSS: free Japanese speech corpus for summarization and simplification
- Utterance-level Sequential Modeling For Deep Gaussian Process Based Speech Synthesis Using Simple Recurrent Unit
- Segment-Level Vectorized Beam Search Based on Partially Autoregressive Inference
- Lifter Training and Sub-band Modeling for Computationally Efficient and High-Quality Voice Conversion Using Spectral Differentials