most citedMusicLM: Generating Music From Text

189 citations · 232 across the 6 of their papers we have counts for

collaborators

6 papers

cs.SD2023

TokenSplit: Using Discrete Speech Representations for Direct, Refined, and Transcript-Conditioned Speech Separation and Recognition

Hakan Erdogan, Scott Wisdom, Xuankai Chang +4

We present TokenSplit, a speech separation model that acts on discrete token sequences. The model is trained on multiple tasks simultaneously: separate and transcribe each speech s…

cs.SD202318 cited

SoundStorm: Efficient Parallel Audio Generation

Zalán Borsos, Matt Sharifi, Damien Vincent +3

We present SoundStorm, a model for efficient, non-autoregressive audio generation. SoundStorm receives as input the semantic tokens of AudioLM, and relies on bidirectional attentio…

cs.SD20231 cited

LMCodec: A Low Bitrate Speech Codec With Causal Transformer Models

Teerapat Jenrungrot, Michael Chinen, W. Bastiaan Kleijn +4

We introduce LMCodec, a causal neural speech codec that provides high quality audio at very low bitrates. The backbone of the system is a causal convolutional codec that encodes au…

cs.SD20233 cited

Speak, Read and Prompt: High-Fidelity Text-to-Speech with Minimal Supervision

Eugene Kharitonov, Damien Vincent, Zalán Borsos +6

We introduce SPEAR-TTS, a multi-speaker text-to-speech (TTS) system that can be trained with minimal supervision. By combining two types of discrete speech representations, we cast…

cs.SD2023189 cited

MusicLM: Generating Music From Text

Andrea Agostinelli, Timo I. Denk, Zalán Borsos +10

We introduce MusicLM, a model generating high-fidelity music from text descriptions such as "a calming violin melody backed by a distorted guitar riff". MusicLM casts the process o…

cs.NI201421 cited

Energy Consumption Of Visual Sensor Networks: Impact Of Spatio-Temporal Coverage

Alessandro Redondi, Dujdow Buranapanichkit, Matteo Cesana +2

Wireless visual sensor networks (VSNs) are expected to play a major role in future IEEE 802.15.4 personal area networks (PAN) under recently-established collision-free medium acces…