activity
20192022
most citedLEAF: A Learnable Frontend for Audio Classification

30 citations · 90 across the 8 of their papers we have counts for

collaborators

16 papers

eess.AS2022

Text-Driven Separation of Arbitrary Sounds

Kevin Kilgour, Beat Gfeller, Qingqing Huang +3

We propose a method of separating a desired sound source from a single-channel mixture, based on either a textual description or a short audio sample of the target source. This is…

cs.SD20222 cited

SpeechPainter: Text-conditioned Speech Inpainting

Zalán Borsos, Matt Sharifi, Marco Tagliasacchi

We propose SpeechPainter, a model for filling in gaps of up to one second in speech samples by leveraging an auxiliary textual input. We demonstrate that the model performs speech…

eess.AS2022

CycleGAN-Based Unpaired Speech Dereverberation

Hannah Muckenhirn, Aleksandr Safin, Hakan Erdogan +4

Typically, neural network-based speech dereverberation models are trained on paired data, composed of a dry utterance and its corresponding reverberant utterance. The main limitati…

cs.LG20211 cited

Data Summarization via Bilevel Optimization

Zalán Borsos, Mojmír Mutný, Marco Tagliasacchi +1

The increasing availability of massive data sets poses a series of challenges for machine learning. Prominent among these is the need to learn models under hardware or human resour…

cs.SD2021

SoundStream: An End-to-End Neural Audio Codec

Neil Zeghidour, Alejandro Luebs, Ahmed Omran +2

We present SoundStream, a novel neural audio codec that can efficiently compress speech, music and general audio at bitrates normally targeted by speech-tailored codecs. SoundStrea…

cs.SD2021

Self-Supervised Learning from Automatically Separated Sound Scenes

Eduardo Fonseca, Aren Jansen, Daniel P. W. Ellis +7

Real-world sound scenes consist of time-varying collections of sound sources, each generating characteristic sound events that are mixed together in audio recordings. The associati…