387 citations · 1.2k across the 18 of their papers we have counts for
Showing 2023 · cs.SDShow all
2 papers · 2 filters
cs.SD2023★ 50 cited
Noise2Music: Text-conditioned Music Generation with Diffusion Models
Qingqing Huang, Daniel S. Park, Tao Wang +12
We introduce Noise2Music, where a series of diffusion models is trained to generate high-quality 30-second music clips from text prompts. Two types of diffusion models, a generator…
cs.SD2023★ 1 cited
Miipher: A Robust Speech Restoration Model Integrating Self-Supervised Speech and Text Representations
Yuma Koizumi, Heiga Zen, Shigeki Karita +7
Speech restoration (SR) is a task of converting degraded speech signals into high-quality ones. In this study, we propose a robust SR model called Miipher, and apply Miipher to a n…