278 citations · 326 across the 5 of their papers we have counts for
17 papers
Learning neural audio features without supervision
Sarthak Yadav, Neil Zeghidour
Deep audio classification, traditionally cast as training a deep neural network on top of mel-filterbanks in a supervised fashion, has recently benefited from two independent lines…
Learning strides in convolutional neural networks
Rachid Riad, Olivier Teboul, David Grangier +1
Convolutional neural networks typically contain several downsampling operators, such as strided convolutions or pooling layers, that progressively reduce the resolution of intermed…
SoundStream: An End-to-End Neural Audio Codec
Neil Zeghidour, Alejandro Luebs, Ahmed Omran +2
We present SoundStream, a novel neural audio codec that can efficiently compress speech, music and general audio at bitrates normally targeted by speech-tailored codecs. SoundStrea…
DIVE: End-to-end Speech Diarization via Iterative Speaker Embedding
Neil Zeghidour, Olivier Teboul, David Grangier
We introduce DIVE, an end-to-end speaker diarization algorithm. Our neural algorithm presents the diarization task as an iterative process: it repeatedly builds a representation fo…
Self-Supervised Learning of Audio Representations from Permutations with Differentiable Ranking
Andrew N Carr, Quentin Berthet, Mathieu Blondel +2
Self-supervised pre-training using so-called "pretext" tasks has recently shown impressive performance across a wide range of modalities. In this work, we advance self-supervised l…
LEAF: A Learnable Frontend for Audio Classification
Neil Zeghidour, Olivier Teboul, Félix de Chaumont Quitry +1
Mel-filterbanks are fixed, engineered audio features which emulate human perception and have been used through the history of audio understanding up to today. However, their undeni…