activity
20172023
most citedMusicLM: Generating Music From Text

189 citations · 300 across the 8 of their papers we have counts for

collaborators
Showing cs.SDShow all

5 papers · 1 filter

cs.SD2023★ 18 cited

SoundStorm: Efficient Parallel Audio Generation

Zalán Borsos, Matt Sharifi, Damien Vincent +3

We present SoundStorm, a model for efficient, non-autoregressive audio generation. SoundStorm receives as input the semantic tokens of AudioLM, and relies on bidirectional attentio…

cs.SD2023★ 3 cited

Speak, Read and Prompt: High-Fidelity Text-to-Speech with Minimal Supervision

Eugene Kharitonov, Damien Vincent, Zalán Borsos +6

We introduce SPEAR-TTS, a multi-speaker text-to-speech (TTS) system that can be trained with minimal supervision. By combining two types of discrete speech representations, we cast…

cs.SD2023★ 189 cited

MusicLM: Generating Music From Text

Andrea Agostinelli, Timo I. Denk, Zalán Borsos +10

We introduce MusicLM, a model generating high-fidelity music from text descriptions such as "a calming violin melody backed by a distorted guitar riff". MusicLM casts the process o…

cs.SD2022★ 2 cited

SpeechPainter: Text-conditioned Speech Inpainting

Zalán Borsos, Matt Sharifi, Marco Tagliasacchi

We propose SpeechPainter, a model for filling in gaps of up to one second in speech samples by leveraging an auxiliary textual input. We demonstrate that the model performs speech…

cs.SD2017★ 21 cited

Now Playing: Continuous low-power music recognition

Blaise Agüera y Arcas, Beat Gfeller, Ruiqi Guo +8

Existing music recognition applications require a connection to a server that performs the actual recognition. In this paper we present a low-power music recognizer that runs entir…