21 citations · 47 across the 4 of their papers we have counts for
6 papers
SpeechPainter: Text-conditioned Speech Inpainting
Zalán Borsos, Matt Sharifi, Marco Tagliasacchi
We propose SpeechPainter, a model for filling in gaps of up to one second in speech samples by leveraging an auxiliary textual input. We demonstrate that the model performs speech…
Training Keyword Spotters with Limited and Synthesized Speech Data
James Lin, Kevin Kilgour, Dominik Roblek +1
With the rise of low power speech-enabled devices, there is a growing demand to quickly produce models for recognizing arbitrary sets of keywords. As with many machine learning tas…
SPICE: Self-supervised Pitch Estimation
Beat Gfeller, Christian Frank, Dominik Roblek +3
We propose a model to estimate the fundamental frequency in monophonic audio, often referred to as pitch estimation. We acknowledge the fact that obtaining ground truth annotations…
Fréchet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms
Kevin Kilgour, Mauricio Zuluaga, Dominik Roblek +1
We propose the Fréchet Audio Distance (FAD), a novel, reference-free evaluation metric for music enhancement algorithms. We demonstrate how typical evaluation metrics for speech en…
Low-Dimensional Bottleneck Features for On-Device Continuous Speech Recognition
David B. Ramsay, Kevin Kilgour, Dominik Roblek +1
Low power digital signal processors (DSPs) typically have a very limited amount of memory in which to cache data. In this paper we develop efficient bottleneck feature (BNF) extrac…
Now Playing: Continuous low-power music recognition
Blaise Agüera y Arcas, Beat Gfeller, Ruiqi Guo +8
Existing music recognition applications require a connection to a server that performs the actual recognition. In this paper we present a low-power music recognizer that runs entir…