activity
20162025
most citedParallel WaveNet: Fast High-Fidelity Speech Synthesis

343 citations · 615 across the 10 of their papers we have counts for

collaborators
Showing cs.SDShow all

7 papers · 1 filter

cs.SD2025

ReverbMiipher: Generative Speech Restoration meets Reverberation Characteristics Controllability

Wataru Nakata, Yuma Koizumi, Shigeki Karita +5

Reverberation encodes spatial information regarding the acoustic source environment, yet traditional Speech Restoration (SR) usually completely removes reverberation. We propose Re…

cs.SD2025

Miipher-2: A Universal Speech Restoration Model for Million-Hour Scale Data Restoration

Shigeki Karita, Yuma Koizumi, Heiga Zen +3

Training data cleaning is a new application for generative model-based speech restoration (SR). This paper introduces Miipher-2, an SR model designed for million-hour scale data, f…

cs.SD20229 cited

Residual Adapters for Few-Shot Text-to-Speech Speaker Adaptation

Nobuyuki Morioka, Heiga Zen, Nanxin Chen +2

Adapting a neural text-to-speech (TTS) model to a target speaker typically involves fine-tuning most if not all of the parameters of a pretrained multi-speaker backbone model. Howe…

cs.SD2021

Parallel Tacotron 2: A Non-Autoregressive Neural TTS Model with Differentiable Duration Modeling

Isaac Elias, Heiga Zen, Jonathan Shen +4

This paper introduces Parallel Tacotron 2, a non-autoregressive neural text-to-speech model with a fully differentiable duration model which does not require supervised duration si…

cs.SD2020

Parallel Tacotron: Non-Autoregressive and Controllable TTS

Isaac Elias, Heiga Zen, Jonathan Shen +4

Although neural end-to-end text-to-speech models can synthesize highly natural speech, there is still room for improvements to its efficiency and naturalness. This paper proposes a…

cs.SD201923 cited

LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Heiga Zen, Viet Dang, Rob Clark +5

This paper introduces a new speech corpus called "LibriTTS" designed for text-to-speech use. It is derived from the original audio and text materials of the LibriSpeech corpus, whi…