Characterizing Speech Adversarial Examples Using Self-Attention U-Net Enhancement
arXiv:2003.13917 · doi:10.1109/ICASSP40776.2020.9053288
Abstract
Recent studies have highlighted adversarial examples as ubiquitous threats to the deep neural network (DNN) based speech recognition systems. In this work, we present a U-Net based attention model, U-Net, to enhance adversarial speech signals. Specifically, we evaluate the model performance by interpretable speech recognition metrics and discuss the model performance by the augmented adversarial training. Our experiments show that our proposed U-Net improves the perceptual evaluation of speech quality (PESQ) from 1.13 to 2.78, speech transmission index (STI) from 0.65 to 0.75, short-term objective intelligibility (STOI) from 0.83 to 0.96 on the task of speech enhancement with adversarial speech examples. We conduct experiments on the automatic speech recognition (ASR) task with adversarial audio attacks. We find that (i) temporal features learned by the attention network are capable of enhancing the robustness of DNN based ASR models; (ii) the generalization power of DNN based ASR model could be enhanced by applying adversarial training with an additive adversarial data augmentation. The ASR metric on word-error-rates (WERs) shows that there is an absolute 2.22 decrease under gradient-based perturbation, and an absolute 2.03 decrease, under evolutionary-optimized perturbation, which suggests that our enhancement models with adversarial training can further secure a resilient ASR system.
The authors have revised some annotations in Table 4 to improve the clarity. The authors thank reading feedbacks from Jonathan Le Roux. The first draft was finished in August 2019. Accepted to IEEE ICASSP 2020
References in corpus (5)
- Explaining and Harnessing Adversarial Examples
- Did you hear that? Adversarial Examples Against Automatic Speech Recognition
- Imperceptible, Robust, and Targeted Adversarial Examples for Automatic Speech Recognition
- Improved Speech Enhancement with the Wave-U-Net
- When Causal Intervention Meets Adversarial Examples and Image Masking for Deep Neural Networks
Cited by in corpus (15)
- Analyzing Upper Bounds on Mean Absolute Errors for Deep Neural Network Based Vector-to-Vector Regression
- Wavelet Channel Attention Module with a Fusion Network for Single Image Deraining
- A Study on Speech Enhancement Based on Diffusion Probabilistic Model
- Decentralizing Feature Extraction with Quantum Convolutional Neural Network for Automatic Speech Recognition
- Voice2Series: Reprogramming Acoustic Models for Time Series Classification
- LAFFNet: A Lightweight Adaptive Feature Fusion Network for Underwater Image Enhancement
- Improved Lite Audio-Visual Speech Enhancement
- PATE-AAE: Incorporating Adversarial Autoencoder into Private Aggregation of Teacher Ensembles for Spoken Command Classification
- Voting for the right answer: Adversarial defense for speaker verification
- Improving Speech Enhancement Performance by Leveraging Contextual Broad Phonetic Class Information
- Neural Kalman Filtering for Speech Enhancement
- Exploring Deep Hybrid Tensor-to-Vector Network Architectures for Regression Based Speech Enhancement
- Incorporating Broad Phonetic Information for Speech Enhancement
- Longer Version for "Deep Context-Encoding Network for Retinal Image Captioning"
- Robust Unsupervised Multi-Object Tracking in Noisy Environments