387 citations · 681 across the 13 of their papers we have counts for
Showing cs.SDShow all
2 papers · 1 filter
cs.SD2023★ 50 cited
Noise2Music: Text-conditioned Music Generation with Diffusion Models
Qingqing Huang, Daniel S. Park, Tao Wang +12
We introduce Noise2Music, where a series of diffusion models is trained to generate high-quality 30-second music clips from text prompts. Two types of diffusion models, a generator…
cs.SD2023
Miipher: A Robust Speech Restoration Model Integrating Self-Supervised Speech and Text Representations
Yuma Koizumi, Heiga Zen, Shigeki Karita +7
Speech restoration (SR) is a task of converting degraded speech signals into high-quality ones. In this study, we propose a robust SR model called Miipher, and apply Miipher to a n…