Interpretable Convolutional Filters with SincNet
arXiv:1811.09725
Abstract
Deep learning is currently playing a crucial role toward higher levels of artificial intelligence. This paradigm allows neural networks to learn complex and abstract representations, that are progressively obtained by combining simpler ones. Nevertheless, the internal "black-box" representations automatically discovered by current neural architectures often suffer from a lack of interpretability, making of primary interest the study of explainable machine learning techniques. This paper summarizes our recent efforts to develop a more interpretable neural model for directly processing speech from the raw waveform. In particular, we propose SincNet, a novel Convolutional Neural Network (CNN) that encourages the first layer to discover more meaningful filters by exploiting parametrized sinc functions. In contrast to standard CNNs, which learn all the elements of each filter, only low and high cutoff frequencies of band-pass filters are directly learned from data. This inductive bias offers a very compact way to derive a customized filter-bank front-end, that only depends on some parameters with a clear physical meaning. Our experiments, conducted on both speaker and speech recognition, show that the proposed architecture converges faster, performs better, and is more interpretable than standard CNNs.
In Proceedings of NIPS@IRASL 2018. arXiv admin note: substantial text overlap with arXiv:1808.00158
References in corpus (7)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
- WaveNet: A Generative Model for Raw Audio
- VoxCeleb: a large-scale speaker identification dataset
- Light Gated Recurrent Units for Speech Recognition
- SampleRNN: An Unconditional End-to-End Neural Audio Generation Model
- End-to-end spoofing detection with raw waveform CLDNNs
Cited by in corpus (15)
- Deep Learning Algorithms for Rotating Machinery Intelligent Diagnosis: An Open Source Benchmark Study
- TFN: An Interpretable Neural Network with Time-Frequency Transform Embedded for Intelligent Fault Diagnosis
- AI-Driven Interface Design for Intelligent Tutoring System Improves Student Engagement
- Interpreting intermediate convolutional layers in unsupervised acoustic word classification
- Short utterance compensation in speaker verification via cosine-based teacher-student learning of speaker embeddings
- An Efficient Explorative Sampling Considering the Generative Boundaries of Deep Generative Neural Networks
- Understanding Semantics from Speech Through Pre-training
- Interpretable Super-Resolution via a Learned Time-Series Representation
- Robust Raw Waveform Speech Recognition Using Relevance Weighted Representations
- Revisiting Representation Learning for Singing Voice Separation with Sinkhorn Distances
- Y-Vector: Multiscale Waveform Encoder for Speaker Embedding
- What does a network layer hear? Analyzing hidden representations of end-to-end ASR through speech synthesis
- Speech recognition for air traffic control via feature learning and end-to-end training
- Learning to fool the speaker recognition
- On Controlled DeEntanglement for Natural Language Processing