activity
20182022
most citedHigh Fidelity Neural Audio Compression

280 citations · 289 across the 5 of their papers we have counts for

collaborators

12 papers

cs.SD20221 cited

Audio Language Modeling using Perceptually-Guided Discrete Representations

Felix Kreuk, Yaniv Taigman, Adam Polyak +4

In this work, we study the task of Audio Language Modeling, in which we aim at learning probabilistic models for audio that can be used for generation and completion. We use a stat…

eess.AS2022280 cited

High Fidelity Neural Audio Compression

Alexandre Défossez, Jade Copet, Gabriel Synnaeve +1

We introduce a state-of-the-art real-time, high-fidelity, audio codec leveraging neural networks. It consists in a streaming encoder-decoder architecture with quantized latent spac…

cs.CL20224 cited

textless-lib: a Library for Textless Spoken Language Processing

Eugene Kharitonov, Jade Copet, Kushal Lakhotia +8

Textless spoken language processing research aims to extend the applicability of standard NLP toolset onto spoken language and languages with few or no textual resources. In this p…

eess.AS20214 cited

fairseq S^2: A Scalable and Integrable Speech Synthesis Toolkit

Changhan Wang, Wei-Ning Hsu, Yossi Adi +5

This paper presents fairseq S^2, a fairseq extension for speech synthesis. We implement a number of autoregressive (AR) and non-AR text-to-speech models, and their multi-speaker va…

cs.SD2021

Speech Resynthesis from Discrete Disentangled Self-Supervised Representations

Adam Polyak, Yossi Adi, Jade Copet +5

We propose using self-supervised discrete representations for the task of speech resynthesis. To generate disentangled representation, we separately extract low-bitrate representat…

cs.CL2021

Generative Spoken Language Modeling from Raw Audio

Kushal Lakhotia, Evgeny Kharitonov, Wei-Ning Hsu +8

We introduce Generative Spoken Language Modeling, the task of learning the acoustic and linguistic characteristics of a language from raw audio (no text, no labels), and a set of m…