most citedNaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models

20 citations · 20 across the 8 of their papers we have counts for

collaborators

8 papers

eess.AS2024

BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec

Detai Xin, Xu Tan, Shinnosuke Takamichi +1

We present BigCodec, a low-bitrate neural speech codec. While recent neural speech codecs have shown impressive progress, their performance significantly deteriorates at low bitrat…

eess.AS202420 cited

NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models

Zeqian Ju, Yuancheng Wang, Kai Shen +16

While recent large-scale text-to-speech (TTS) models have achieved significant progress, they still fall short in speech quality, similarity, and prosody. Considering speech intric…

cs.SD2024

Building speech corpus with diverse voice characteristics for its prompt-based representation

Aya Watanabe, Shinnosuke Takamichi, Yuki Saito +3

In text-to-speech synthesis, the ability to control voice characteristics is vital for various applications. By leveraging thriving text prompt-based generation techniques, it shou…

cs.SD2023

Coco-Nut: Corpus of Japanese Utterance and Voice Characteristics Description for Prompt-based Control

Aya Watanabe, Shinnosuke Takamichi, Yuki Saito +3

In text-to-speech, controlling voice characteristics is important in achieving various-purpose speech synthesis. Considering the success of text-conditioned generation, such as tex…

cs.CL2023

How Generative Spoken Language Modeling Encodes Noisy Speech: Investigation from Phonetics to Syntactics

Joonyong Park, Shinnosuke Takamichi, Tomohiko Nakamura +3

We examine the speech modeling potential of generative spoken language modeling (GSLM), which involves using learned symbols derived from data rather than phonemes for speech analy…

cs.SD2023

Laughter Synthesis using Pseudo Phonetic Tokens with a Large-scale In-the-wild Laughter Corpus

Detai Xin, Shinnosuke Takamichi, Ai Morimatsu +1

We present a large-scale in-the-wild Japanese laughter corpus and a laughter synthesis method. Previous work on laughter synthesis lacks not only data but also proper ways to repre…