Low-frequency Compensated Synthetic Impulse Responses for Improved Far-field Speech Recognition
arXiv:1910.10815 · doi:10.1109/ICASSP40776.2020.9054454
Abstract
We propose a method for generating low-frequency compensated synthetic impulse responses that improve the performance of far-field speech recognition systems trained on artificially augmented datasets. We design linear-phase filters that adapt the simulated impulse responses to equalization distributions corresponding to real-world captured impulse responses. Our filtered synthetic impulse responses are then used to augment clean speech data from LibriSpeech dataset [1]. We evaluate the performance of our method on the real-world LibriSpeech test set. In practice, our low-frequency compensated synthetic dataset can reduce the word-error-rate by up to 8.8% for far-field speech recognition.
Accepted to ICASSP 2020
References in corpus (5)
- Building and Evaluation of a Real Room Impulse Response Dataset
- Regression and Classification for Direction-of-Arrival Estimation with Convolutional Recurrent Neural Networks
- Scene-Aware Audio Rendering via Deep Acoustic Analysis
- Improving Reverberant Speech Training Using Diffuse Acoustic Simulation
- Interactive Sound Rendering on Mobile Devices using Ray-Parameterized Reverberation Filters
Cited by in corpus (5)
- Scene-Aware Audio Rendering via Deep Acoustic Analysis
- Improving Reverberant Speech Training Using Diffuse Acoustic Simulation
- Sound Synthesis, Propagation, and Rendering: A Survey
- Scene-aware Far-field Automatic Speech Recognition
- A study on more realistic room simulation for far-field keyword spotting