activity
20172026
most citedLow Bit-Rate Speech Coding with VQ-VAE and a WaveNet Decoder

120 citations · 146 across the 4 of their papers we have counts for

collaborators

7 papers

eess.AS2026

PoDAR: Power-Disentangled Audio Representation for Generative Modeling

Alejandro Luebs, Mithilesh Vaidya, Ishaan Kumar +5

The performance of audio latent diffusion models is primarily governed by generator expressivity and the modelability of the underlying latent space. While recent research has focu…

cs.SD2023★ 24 cited

High-Fidelity Audio Compression with Improved RVQGAN

Rithesh Kumar, Prem Seetharaman, Alejandro Luebs +2

Language models have been successfully used to model natural signals, such as images, speech, and music. A key component of these models is a high quality neural compression model…

cs.SD2021

SoundStream: An End-to-End Neural Audio Codec

Neil Zeghidour, Alejandro Luebs, Ahmed Omran +2

We present SoundStream, a novel neural audio codec that can efficiently compress speech, music and general audio at bitrates normally targeted by speech-tailored codecs. SoundStrea…

eess.AS2021

Handling Background Noise in Neural Speech Generation

Tom Denton, Alejandro Luebs, Felicia S. C. Lim +4

Recent advances in neural-network based generative modeling of speech has shown great potential for speech coding. However, the performance of such models drops when the input is n…

eess.AS2021

Generative Speech Coding with Predictive Variance Regularization

W. Bastiaan Kleijn, Andrew Storus, Michael Chinen +5

The recent emergence of machine-learning based generative models for speech suggests a significant reduction in bit rate for speech codecs is possible. However, the performance of…

cs.LG2019★ 120 cited

Low Bit-Rate Speech Coding with VQ-VAE and a WaveNet Decoder

Cristina Gârbacea, Aäron van den Oord, Yazhe Li +4

In order to efficiently transmit and store speech signals, speech codecs create a minimally redundant representation of the input signal which is then decoded at the receiver with…