189 citations · 227 across the 3 of their papers we have counts for
3 papers
Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
Philip Anastassiou, Jiawei Chen, Jitong Chen +43
We introduce Seed-TTS, a family of large-scale autoregressive text-to-speech (TTS) models capable of generating speech that is virtually indistinguishable from human speech. Seed-T…
MusicLM: Generating Music From Text
Andrea Agostinelli, Timo I. Denk, Zalán Borsos +10
We introduce MusicLM, a model generating high-fidelity music from text descriptions such as "a calming violin melody backed by a distorted guitar riff". MusicLM casts the process o…
MuLan: A Joint Embedding of Music Audio and Natural Language
Qingqing Huang, Aren Jansen, Joonseok Lee +3
Music tagging and content-based retrieval systems have traditionally been constructed using pre-defined ontologies covering a rigid set of music attributes or text queries. This pa…