280 citations · 289 across the 5 of their papers we have counts for
12 papers
Audio Language Modeling using Perceptually-Guided Discrete Representations
Felix Kreuk, Yaniv Taigman, Adam Polyak +4
In this work, we study the task of Audio Language Modeling, in which we aim at learning probabilistic models for audio that can be used for generation and completion. We use a stat…
High Fidelity Neural Audio Compression
Alexandre Défossez, Jade Copet, Gabriel Synnaeve +1
We introduce a state-of-the-art real-time, high-fidelity, audio codec leveraging neural networks. It consists in a streaming encoder-decoder architecture with quantized latent spac…
textless-lib: a Library for Textless Spoken Language Processing
Eugene Kharitonov, Jade Copet, Kushal Lakhotia +8
Textless spoken language processing research aims to extend the applicability of standard NLP toolset onto spoken language and languages with few or no textual resources. In this p…
fairseq S^2: A Scalable and Integrable Speech Synthesis Toolkit
Changhan Wang, Wei-Ning Hsu, Yossi Adi +5
This paper presents fairseq S^2, a fairseq extension for speech synthesis. We implement a number of autoregressive (AR) and non-AR text-to-speech models, and their multi-speaker va…
Speech Resynthesis from Discrete Disentangled Self-Supervised Representations
Adam Polyak, Yossi Adi, Jade Copet +5
We propose using self-supervised discrete representations for the task of speech resynthesis. To generate disentangled representation, we separately extract low-bitrate representat…
Generative Spoken Language Modeling from Raw Audio
Kushal Lakhotia, Evgeny Kharitonov, Wei-Ning Hsu +8
We introduce Generative Spoken Language Modeling, the task of learning the acoustic and linguistic characteristics of a language from raw audio (no text, no labels), and a set of m…