50 citations · 50 across the 1 of their papers we have counts for
2 papers
cs.CL2023
AudioPaLM: A Large Language Model That Can Speak and Listen
Paul K. Rubenstein, Chulayuth Asawaroengchai, Duc Dung Nguyen +27
We introduce AudioPaLM, a large language model for speech understanding and generation. AudioPaLM fuses text-based and speech-based language models, PaLM-2 [Anil et al., 2023] and…
cs.SD2023★ 50 cited
Noise2Music: Text-conditioned Music Generation with Diffusion Models
Qingqing Huang, Daniel S. Park, Tao Wang +12
We introduce Noise2Music, where a series of diffusion models is trained to generate high-quality 30-second music clips from text prompts. Two types of diffusion models, a generator…