30 citations · 90 across the 8 of their papers we have counts for
16 papers
Text-Driven Separation of Arbitrary Sounds
Kevin Kilgour, Beat Gfeller, Qingqing Huang +3
We propose a method of separating a desired sound source from a single-channel mixture, based on either a textual description or a short audio sample of the target source. This is…
SpeechPainter: Text-conditioned Speech Inpainting
Zalán Borsos, Matt Sharifi, Marco Tagliasacchi
We propose SpeechPainter, a model for filling in gaps of up to one second in speech samples by leveraging an auxiliary textual input. We demonstrate that the model performs speech…
CycleGAN-Based Unpaired Speech Dereverberation
Hannah Muckenhirn, Aleksandr Safin, Hakan Erdogan +4
Typically, neural network-based speech dereverberation models are trained on paired data, composed of a dry utterance and its corresponding reverberant utterance. The main limitati…
Data Summarization via Bilevel Optimization
Zalán Borsos, Mojmír Mutný, Marco Tagliasacchi +1
The increasing availability of massive data sets poses a series of challenges for machine learning. Prominent among these is the need to learn models under hardware or human resour…
SoundStream: An End-to-End Neural Audio Codec
Neil Zeghidour, Alejandro Luebs, Ahmed Omran +2
We present SoundStream, a novel neural audio codec that can efficiently compress speech, music and general audio at bitrates normally targeted by speech-tailored codecs. SoundStrea…
Self-Supervised Learning from Automatically Separated Sound Scenes
Eduardo Fonseca, Aren Jansen, Daniel P. W. Ellis +7
Real-world sound scenes consist of time-varying collections of sound sources, each generating characteristic sound events that are mixed together in audio recordings. The associati…