21 citations · 22 across the 4 of their papers we have counts for
4 papers
AV-TranSpeech: Audio-Visual Robust Speech-to-Speech Translation
Rongjie Huang, Huadai Liu, Xize Cheng +8
Direct speech-to-speech translation (S2ST) aims to convert speech from one language into another, and has demonstrated significant progress to date. Despite the recent success, cur…
Wav2SQL: Direct Generalizable Speech-To-SQL Parsing
Huadai Liu, Rongjie Huang, Jinzheng He +4
Speech-to-SQL (S2SQL) aims to convert spoken questions into SQL queries given relational databases, which has been traditionally implemented in a cascaded manner while facing the f…
RMSSinger: Realistic-Music-Score based Singing Voice Synthesis
Jinzheng He, Jinglin Liu, Zhenhui Ye +4
We are interested in a challenging task, Realistic-Music-Score based Singing Voice Synthesis (RMS-SVS). RMS-SVS aims to generate high-quality singing voices given realistic music s…
ProDiff: Progressive Fast Diffusion Model For High-Quality Text-to-Speech
Rongjie Huang, Zhou Zhao, Huadai Liu +3
Denoising diffusion probabilistic models (DDPMs) have recently achieved leading performances in many generative tasks. However, the inherited iterative sampling process costs hinde…