I'm Sorry for Your Loss: Spectrally-Based Audio Distances Are Bad at Pitch
arXiv:2012.04572
Abstract
Growing research demonstrates that synthetic failure modes imply poor generalization. We compare commonly used audio-to-audio losses on a synthetic benchmark, measuring the pitch distance between two stationary sinusoids. The results are surprising: many have poor sense of pitch direction. These shortcomings are exposed using simple rank assumptions. Our task is trivial for humans but difficult for these audio distances, suggesting significant progress can be made in self-supervised audio learning by improving current losses.
ICBINB@NeurIPS 2020
References in corpus (12)
- Distilling the Knowledge in a Neural Network
- Natural Language Processing (almost) from Scratch
- WaveNet: A Generative Model for Raw Audio
- An Overview of Multi-Task Learning in Deep Neural Networks
- MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis
- Underspecification Presents Challenges for Credibility in Modern Machine Learning
- SampleRNN: An Unconditional End-to-End Neural Audio Generation Model
- Low Bit-Rate Speech Coding with VQ-VAE and a WaveNet Decoder
- Jukebox: A Generative Model for Music
- DDSP: Differentiable Digital Signal Processing
- Understanding the Failure Modes of Out-of-Distribution Generalization
- Quantifying the Preferential Direction of the Model Gradient in Adversarial Training With Projected Gradient Descent