189 citations · 267 across the 13 of their papers we have counts for
9 papers · 1 filter
SpectroStream: A Versatile Neural Codec for General Audio
Yunpeng Li, Kehang Han, Brian McWilliams +2
We propose SpectroStream, a full-band multi-channel neural audio codec. Successor to the well-established SoundStream, SpectroStream extends its capability beyond 24 kHz monophonic…
Live Music Models
Lyria Team, Antoine Caillon, Brian McWilliams +33
We introduce a new class of generative models for music called live music models that produce a continuous stream of music in real-time with synchronized user control. We release M…
TokenSplit: Using Discrete Speech Representations for Direct, Refined, and Transcript-Conditioned Speech Separation and Recognition
Hakan Erdogan, Scott Wisdom, Xuankai Chang +4
We present TokenSplit, a speech separation model that acts on discrete token sequences. The model is trained on multiple tasks simultaneously: separate and transcribe each speech s…
SoundStorm: Efficient Parallel Audio Generation
Zalán Borsos, Matt Sharifi, Damien Vincent +3
We present SoundStorm, a model for efficient, non-autoregressive audio generation. SoundStorm receives as input the semantic tokens of AudioLM, and relies on bidirectional attentio…
LMCodec: A Low Bitrate Speech Codec With Causal Transformer Models
Teerapat Jenrungrot, Michael Chinen, W. Bastiaan Kleijn +4
We introduce LMCodec, a causal neural speech codec that provides high quality audio at very low bitrates. The backbone of the system is a causal convolutional codec that encodes au…
Speak, Read and Prompt: High-Fidelity Text-to-Speech with Minimal Supervision
Eugene Kharitonov, Damien Vincent, Zalán Borsos +6
We introduce SPEAR-TTS, a multi-speaker text-to-speech (TTS) system that can be trained with minimal supervision. By combining two types of discrete speech representations, we cast…