Universal Paralinguistic Speech Representations Using Self-Supervised Conformers
arXiv:2110.04621 · doi:10.1109/ICASSP43922.2022.9747197
Abstract
Many speech applications require understanding aspects beyond the words being spoken, such as recognizing emotion, detecting whether the speaker is wearing a mask, or distinguishing real from synthetic speech. In this work, we introduce a new state-of-the-art paralinguistic representation derived from large-scale, fully self-supervised training of a 600M+ parameter Conformer-based architecture. We benchmark on a diverse set of speech tasks and demonstrate that simple linear classifiers trained on top of our time-averaged representation outperform nearly all previous results, in some cases by large margins. Our analyses of context-window size demonstrate that, surprisingly, 2 second context-windows achieve 96\% the performance of the Conformers that use the full long-term context on 7 out of 9 tasks. Furthermore, while the best per-task representations are extracted internally in the network, stable performance across several layers allows a single universal representation to reach near optimal performance on all tasks.
References in corpus (3)
Cited by in corpus (6)
- TRILLsson: Distilled Universal Paralinguistic Speech Representations
- Designing and Evaluating Speech Emotion Recognition Systems: A reality check case study with IEMOCAP
- LanSER: Language-Model Supported Speech Emotion Recognition
- Zero-shot text-to-speech synthesis conditioned using self-supervised speech representation model
- Clinical BERTScore: An Improved Measure of Automatic Speech Recognition Performance in Clinical Settings
- Face and Voice Cross-modal Association with Learning Convex Feature Embedding