Deep Scattering Spectrum
arXiv:1304.6763 · doi:10.1109/TSP.2014.2326991
Abstract
A scattering transform defines a locally translation invariant representation which is stable to time-warping deformations. It extends MFCC representations by computing modulation spectrum coefficients of multiple orders, through cascades of wavelet convolutions and modulus operators. Second-order scattering coefficients characterize transient phenomena such as attacks and amplitude modulation. A frequency transposition invariant representation is obtained by applying a scattering transform along log-frequency. State-the-of-art classification results are obtained for musical genre and phone classification on GTZAN and TIMIT databases, respectively.
Cited by in corpus (87)
- Understanding Deep Convolutional Networks
- A new approach to observational cosmology using the scattering transform
- Fully Learnable Deep Wavelet Transform for Unsupervised Monitoring of High-Frequency Time Series
- The GTZAN dataset: Its contents, its faults, their effects on evaluation, and its future use
- SPICE: Self-supervised Pitch Estimation
- Wavelet Scattering Regression of Quantum Chemical Energies
- Joint Time-Frequency Scattering
- Modulation spectral features for speech emotion recognition using deep neural networks
- Intermittent process analysis with scattering moments
- Integrating the Data Augmentation Scheme with Various Classifiers for Acoustic Scene Modeling
- Differentially Private Learning Needs Better Features (or Much More Data)
- Weak lensing scattering transform: dark energy and neutrino mass sensitivity
- Going Beyond the Galaxy Power Spectrum: an Analysis of BOSS Data with Wavelet Scattering Transforms
- Demucs: Deep Extractor for Music Sources with extra unlabeled data remixed
- Investigation of Data Augmentation Techniques for Disordered Speech Recognition
- Joint Time-Frequency Scattering for Audio Classification
- Precise Cosmological Constraints from BOSS Galaxy Clustering with a Simulation-Based Emulator of the Wavelet Scattering Transform
- Audio Texture Synthesis with Scattering Moments
- Radar-based Materials Classification Using Deep Wavelet Scattering Transform: A Comparison of Centimeter vs. Millimeter Wave Units
- LEAF: A Learnable Frontend for Audio Classification
- Wavelet Moments for Cosmological Parameter Estimation
- A Deep Representation for Invariance And Music Classification
- How informative are summaries of the cosmic 21-cm signal?
- Masked Conditional Neural Networks for Audio Classification
- Statistical Hypothesis Testing Based on Machine Learning: Large Deviations Analysis
- Physics-assisted Generative Adversarial Network for X-Ray Tomography
- Wavelet Scattering on the Pitch Spiral
- Discriminative Segmental Cascades for Feature-Rich Phone Recognition
- Analysis of constant-Q filterbank based representations for speech emotion recognition
- Geometric Scattering Attention Networks
- Discrete Deep Feature Extraction: A Theory and New Architectures
- Towards unveiling the large-scale nature of gravity with the wavelet scattering transform
- Combining Scatter Transform and Deep Neural Networks for Multilabel Electrocardiogram Signal Classification
- Learning Waveform-Based Acoustic Models using Deep Variational Convolutional Neural Networks
- Exploring spectro-temporal features in end-to-end convolutional neural networks
- Sparse Pursuit and Dictionary Learning for Blind Source Separation in Polyphonic Music Recordings
- Geometric Scattering for Graph Data Analysis
- Scattering Transform Based Image Clustering using Projection onto Orthogonal Complement
- SigWavNet: Learning Multiresolution Signal Wavelet Network for Speech Emotion Recognition
- Fatigue monitoring and maneuver identification for vehicle fleets using a virtual sensing approach
- Rotationally invariant time-frequency scattering transforms
- Geometric Scattering on Manifolds
- Degrees of Freedom Analysis of Unrolled Neural Networks
- ACGAN-based Data Augmentation Integrated with Long-term Scalogram for Acoustic Scene Classification
- Current state of nonlinear-type time-frequency analysis and applications to high-frequency biomedical signals
- Automated Polysomnography Analysis for Detection of Non-Apneic and Non-Hypopneic Arousals using Feature Engineering and a Bidirectional LSTM Network
- Gabor frames and deep scattering networks in audio processing
- Geometric Wavelet Scattering Networks on Compact Riemannian Manifolds
- Improving Machine Hearing on Limited Data Sets
- Adaptive DCTNet for Audio Signal Classification
- Sequence Prediction with Neural Segmental Models
- Feedback-Controlled Sequential Lasso Screening
- Texture Retrieval via the Scattering Transform
- The Shape of RemiXXXes to Come: Audio Texture Synthesis with Time-frequency Scattering
- Complementing Handcrafted Features with Raw Waveform Using a Light-weight Auxiliary Model
- Learnable MFCCs for Speaker Verification
- Clustering Noisy Signals with Structured Sparsity Using Time-Frequency Representation
- DCTNet and PCANet for acoustic signal feature extraction
- Fast Chirplet Transform to Enhance CNN Machine Listening - Validation on Animal calls and Speech
- wav2shape: Hearing the Shape of a Drum Machine
- R3Net: Random Weights, Rectifier Linear Units and Robustness for Artificial Neural Network
- Diffuse to fuse EEG spectra -- intrinsic geometry of sleep dynamics for classification
- Learning An Invariant Speech Representation
- Energy Propagation in Deep Convolutional Neural Networks
- Tensor Switching Networks
- An Overview of Lead and Accompaniment Separation in Music
- Robust Unsupervised Transient Detection With Invariant Representation based on the Scattering Network
- Deep scattering transform applied to note onset detection and instrument recognition
- Continuous Generative Neural Networks: A Wavelet-Based Architecture in Function Spaces
- Y-Vector: Multiscale Waveform Encoder for Speaker Embedding
- Learning Boolean functions with concentrated spectra
- Learning to detect dysarthria from raw speech
- One or Two Components? The Scattering Transform Answers
- Phase-Based Signal Representations for Scattering
- Three-Dimensional Fourier Scattering Transform and Classification of Hyperspectral Images
- A Methodology for Exploring Deep Convolutional Features in Relation to Hand-Crafted Features with an Application to Music Audio Modeling
- Classifier-independent Lower-Bounds for Adversarial Robustness
- Central and Non-central Limit Theorems arising from the Scattering Transform and its Neural Activation Generalization
- Interpretable Image Clustering via Diffeomorphism-Aware K-Means
- Une ou deux composantes ? La réponse de la diffusion en ondelettes
- Graceful Forgetting II. Data as a Process
- Wavelet Classification for Over-the-Air Non-Orthogonal Waveforms
- Deep Autoencoders: From Understanding to Generalization Guarantees
- MS-SincResNet: Joint learning of 1D and 2D kernels using multi-scale SincNet and ResNet for music genre classification
- Transformée en scattering sur la spirale temps-chroma-octave
- Dynamic Texture Recognition via Nuclear Distances on Kernelized Scattering Histogram Spaces
- Scattering Features for Multimodal Gait Recognition