Word-level Embeddings for Cross-Task Transfer Learning in Speech Processing
arXiv:1910.09909 · doi:10.23919/EUSIPCO54536.2021.9616254
Abstract
Recent breakthroughs in deep learning often rely on representation learning and knowledge transfer. In recent years, unsupervised and self-supervised techniques for learning speech representation were developed to foster automatic speech recognition. Up to date, most of these approaches are task-specific and designed for within-task transfer learning between different datasets or setups of a particular task. In turn, learning task-independent representation of speech and cross-task applications of transfer learning remain less common. Here, we introduce an encoder capturing word-level representations of speech for cross-task transfer learning. We demonstrate the application of the pre-trained encoder in four distinct speech and audio processing tasks: (i) speech enhancement, (ii) language identification, (iii) speech, noise, and music classification, and (iv) speaker identification. In each task, we compare the performance of our cross-task transfer learning approach to task-specific baselines. Our results show that the speech representation captured by the encoder through the pre-training is transferable across distinct speech processing tasks and datasets. Notably, even simple applications of our pre-trained encoder outperformed task-specific methods, or were comparable, depending on the task.
Published at EUSIPCO 2021
References in corpus (7)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition
- vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations
- Unsupervised speech representation learning using WaveNet autoencoders
- Deep speech inpainting of time-frequency masks
- Self-supervised audio representation learning for mobile devices
- Learning audio representations via phase prediction