activity
20162023
most citedAudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head

21 citations · 109 across the 27 of their papers we have counts for

collaborators

27 papers

cs.CL20232 cited

MSSRNet: Manipulating Sequential Style Representation for Unsupervised Text Style Transfer

Yazheng Yang, Zhou Zhao, Qi Liu

Unsupervised text style transfer task aims to rewrite a text into target style while preserving its main content. Traditional methods rely on the use of a fixed-sized vector to reg…

cs.CL2023

OpenSR: Open-Modality Speech Recognition via Maintaining Multi-Modality Alignment

Xize Cheng, Tao Jin, Linjun Li +3

Speech Recognition builds a bridge between the multimedia streaming (audio-only, visual-only or audio-visual) and the corresponding text transcription. However, when training the s…

eess.AS202316 cited

Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias

Ziyue Jiang, Yi Ren, Zhenhui Ye +9

Scaling text-to-speech to a large and wild dataset has been proven to be highly effective in achieving timbre and speech style generalization, particularly in zero-shot TTS. Howeve…

cs.CV20232 cited

Detector Guidance for Multi-Object Text-to-Image Generation

Luping Liu, Zijian Zhang, Yi Ren +3

Diffusion models have demonstrated impressive performance in text-to-image generation. They utilize a text encoder and cross-attention blocks to infuse textual information into ima…

eess.AS20232 cited

Make-A-Voice: Unified Voice Synthesis With Discrete Representation

Rongjie Huang, Chunlei Zhang, Yongqi Wang +7

Various applications of voice synthesis have been developed independently despite the fact that they generate "voice" as output in common. In addition, the majority of voice synthe…

cs.SD202310 cited

Make-An-Audio 2: Temporal-Enhanced Text-to-Audio Generation

Jiawei Huang, Yi Ren, Rongjie Huang +7

Large diffusion models have been successful in text-to-audio (T2A) synthesis tasks, but they often suffer from common issues such as semantic misalignment and poor temporal consist…