activity
20202024
most citedMake-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion Models

47 citations · 162 across the 25 of their papers we have counts for

collaborators
Showing cs.CLShow all

6 papers · 1 filter

cs.CL2023

AV-TranSpeech: Audio-Visual Robust Speech-to-Speech Translation

Rongjie Huang, Huadai Liu, Xize Cheng +8

Direct speech-to-speech translation (S2ST) aims to convert speech from one language into another, and has demonstrated significant progress to date. Despite the recent success, cur…

cs.CL2023★ 21 cited

AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head

Rongjie Huang, Mingze Li, Dongchao Yang +10

Large language models (LLMs) have exhibited remarkable capabilities across a variety of domains and tasks, challenging our understanding of learning and cognition. Despite the rece…

cs.CL2023★ 1 cited

MUG: A General Meeting Understanding and Generation Benchmark

Qinglin Zhang, Chong Deng, Jiaqing Liu +7

Listening to long video/audio recordings from video conferencing and online courses for acquiring information is extremely inefficient. Even after ASR systems transcribe recordings…

cs.CL2023

Overview of the ICASSP 2023 General Meeting Understanding and Generation Challenge (MUG)

Qinglin Zhang, Chong Deng, Jiaqing Liu +7

ICASSP2023 General Meeting Understanding and Generation Challenge (MUG) focuses on prompting a wide range of spoken language processing (SLP) research on meeting transcripts, as SL…

cs.CL2021

EMOVIE: A Mandarin Emotion Speech Dataset with a Simple Emotional Text-to-Speech Model

Chenye Cui, Yi Ren, Jinglin Liu +4

Recently, there has been an increasing interest in neural speech synthesis. While the deep neural network achieves the state-of-the-art result in text-to-speech (TTS) tasks, how to…

cs.CL2020★ 1 cited

A Study of Non-autoregressive Model for Sequence Generation

Yi Ren, Jinglin Liu, Xu Tan +3

Non-autoregressive (NAR) models generate all the tokens of a sequence in parallel, resulting in faster generation speed compared to their autoregressive (AR) counterparts but at th…