activity
20172020
most citedSelf-supervised audio representation learning for mobile devices

27 citations · 103 across the 7 of their papers we have counts for

collaborators

12 papers

eess.AS2020

One-shot conditional audio filtering of arbitrary sounds

Beat Gfeller, Dominik Roblek, Marco Tagliasacchi

We consider the problem of separating a particular sound source from a single-channel mixture, based on only a short sample of the target source. Using SoundFilter, a wave-to-wave…

eess.AS2020

Real-time Speech Frequency Bandwidth Extension

Yunpeng Li, Marco Tagliasacchi, Oleg Rybakov +2

In this paper we propose a lightweight model for frequency bandwidth extension of speech signals, increasing the sampling frequency from 8kHz to 16kHz while restoring the high freq…

eess.AS2020

SEANet: A Multi-modal Speech Enhancement Network

Marco Tagliasacchi, Yunpeng Li, Karolis Misiunas +1

We explore the possibility of leveraging accelerometer data to perform speech enhancement in very noisy conditions. Although it is possible to only partially reconstruct user's spe…

eess.AS20207 cited

Training Keyword Spotters with Limited and Synthesized Speech Data

James Lin, Kevin Kilgour, Dominik Roblek +1

With the rise of low power speech-enabled devices, there is a growing demand to quickly produce models for recognizing arbitrary sets of keywords. As with many machine learning tas…

eess.AS201910 cited

Learning audio representations via phase prediction

Félix de Chaumont Quitry, Marco Tagliasacchi, Dominik Roblek

We learn audio representations by solving a novel self-supervised learning task, which consists of predicting the phase of the short-time Fourier transform from its magnitude. A co…

eess.AS2019

SPICE: Self-supervised Pitch Estimation

Beat Gfeller, Christian Frank, Dominik Roblek +3

We propose a model to estimate the fundamental frequency in monophonic audio, often referred to as pitch estimation. We acknowledge the fact that obtaining ground truth annotations…