activity
20222026
most citedNaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models

20 citations · 20 across the 12 of their papers we have counts for

collaborators
Showing cs.SDShow all

9 papers · 1 filter

cs.SD2026

LongCat-AudioDiT: High-Fidelity Diffusion Text-to-Speech in the Waveform Latent Space

Detai Xin, Shujie Hu, Chengzuo Yang +4

We present LongCat-AudioDiT, a novel, non-autoregressive diffusion-based text-to-speech (TTS) model that achieves state-of-the-art (SOTA) performance. Unlike previous methods that…

cs.SD2024

Building speech corpus with diverse voice characteristics for its prompt-based representation

Aya Watanabe, Shinnosuke Takamichi, Yuki Saito +3

In text-to-speech synthesis, the ability to control voice characteristics is vital for various applications. By leveraging thriving text prompt-based generation techniques, it shou…

cs.SD2023

JVNV: A Corpus of Japanese Emotional Speech with Verbal Content and Nonverbal Expressions

Detai Xin, Junfeng Jiang, Shinnosuke Takamichi +3

We present the JVNV, a Japanese emotional speech corpus with verbal content and nonverbal vocalizations whose scripts are generated by a large-scale language model. Existing emotio…

cs.SD2023

Coco-Nut: Corpus of Japanese Utterance and Voice Characteristics Description for Prompt-based Control

Aya Watanabe, Shinnosuke Takamichi, Yuki Saito +3

In text-to-speech, controlling voice characteristics is important in achieving various-purpose speech synthesis. Considering the success of text-conditioned generation, such as tex…

cs.SD2023

Laughter Synthesis using Pseudo Phonetic Tokens with a Large-scale In-the-wild Laughter Corpus

Detai Xin, Shinnosuke Takamichi, Ai Morimatsu +1

We present a large-scale in-the-wild Japanese laughter corpus and a laughter synthesis method. Previous work on laughter synthesis lacks not only data but also proper ways to repre…

cs.SD2023

JNV Corpus: A Corpus of Japanese Nonverbal Vocalizations with Diverse Phrases and Emotions

Detai Xin, Shinnosuke Takamichi, Hiroshi Saruwatari

We present JNV (Japanese Nonverbal Vocalizations) corpus, a corpus of Japanese nonverbal vocalizations (NVs) with diverse phrases and emotions. Existing Japanese NV corpora lack ph…