Probing Speech Emotion Recognition Transformers for Linguistic Knowledge
arXiv:2204.00400 · doi:10.21437/Interspeech.2022-10371
Abstract
Large, pre-trained neural networks consisting of self-attention layers (transformers) have recently achieved state-of-the-art results on several speech emotion recognition (SER) datasets. These models are typically pre-trained in self-supervised manner with the goal to improve automatic speech recognition performance -- and thus, to understand linguistic information. In this work, we investigate the extent in which this information is exploited during SER fine-tuning. Using a reproducible methodology based on open-source tools, we synthesise prosodically neutral speech utterances while varying the sentiment of the text. Valence predictions of the transformer model are very reactive to positive and negative sentiment content, as well as negations, but not to intensifiers or reducers, while none of those linguistic features impact arousal or dominance. These findings show that transformers can successfully leverage linguistic information to improve their valence predictions, and that linguistic analysis should be included in their testing.
Accepted in INTERSPEECH 2022
References in corpus (5)
- On the Opportunities and Risks of Foundation Models
- Emotion Intensity and its Control for Emotional Voice Conversion
- A Fine-tuned Wav2vec 2.0/HuBERT Benchmark For Speech Emotion Recognition, Speaker Verification and Spoken Language Understanding
- Normalise for Fairness: A Simple Normalisation Technique for Fairness in Regression Machine Learning Problems
- Best Practices for Noise-Based Augmentation to Improve the Performance of Deployable Speech-Based Emotion Recognition Systems
Cited by in corpus (5)
- Dawn of the transformer era in speech emotion recognition: closing the valence gap
- An Overview of Affective Speech Synthesis and Conversion in the Deep Learning Era
- Estimating the Uncertainty in Emotion Attributes using Deep Evidential Regression
- Leveraging Semantic Information for Efficient Self-Supervised Emotion Recognition with Audio-Textual Distilled Models
- Abusive Speech Detection in Indic Languages Using Acoustic Features