most citedDiffsound: Discrete Diffusion Model for Text-to-sound Generation

9 citations · 15 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CL2023

Bayes Risk Transducer: Transducer with Controllable Alignment Prediction

Jinchuan Tian, Jianwei Yu, Hangting Chen +4

Automatic speech recognition (ASR) based on transducers is widely used. In training, a transducer maximizes the summed posteriors of all paths. The path with the highest posterior…

eess.AS2023★ 2 cited

Make-A-Voice: Unified Voice Synthesis With Discrete Representation

Rongjie Huang, Chunlei Zhang, Yongqi Wang +7

Various applications of voice synthesis have been developed independently despite the fact that they generate "voice" as output in common. In addition, the majority of voice synthe…

cs.CL2022★ 4 cited

Bayes risk CTC: Controllable CTC alignment in Sequence-to-Sequence tasks

Jinchuan Tian, Brian Yan, Jianwei Yu +3

Sequence-to-Sequence (seq2seq) tasks transcribe the input sequence to a target sequence. The Connectionist Temporal Classification (CTC) criterion is widely used in multiple seq2se…

cs.SD2022★ 9 cited

Diffsound: Discrete Diffusion Model for Text-to-sound Generation

Dongchao Yang, Jianwei Yu, Helin Wang +4

Generating sound effects that humans want is an important topic. However, there are few studies in this area for sound generation. In this study, we investigate generating sound co…

cs.SD2022

Automatic Prosody Annotation with Pre-Trained Text-Speech Model

Ziqian Dai, Jianwei Yu, Yan Wang +5

Prosodic boundary plays an important role in text-to-speech synthesis (TTS) in terms of naturalness and readability. However, the acquisition of prosodic boundary labels relies on…