What all do audio transformer models hear? Probing Acoustic Representations for Language Delivery and its Structure
arXiv:2101.00387
Abstract
In recent times, BERT based transformer models have become an inseparable part of the 'tech stack' of text processing models. Similar progress is being observed in the speech domain with a multitude of models observing state-of-the-art results by using audio transformer models to encode speech. This begs the question of what are these audio transformer models learning. Moreover, although the standard methodology is to choose the last layer embedding for any downstream task, but is it the optimal choice? We try to answer these questions for the two recent audio transformer models, Mockingjay and wave2vec2.0. We compare them on a comprehensive set of language delivery and structure features including audio, fluency and pronunciation features. Additionally, we probe the audio models' understanding of textual surface, syntax, and semantic features and compare them to BERT. We do this over exhaustive settings for native, non-native, synthetic, read and spontaneous speech datasets
References in corpus (9)
- Mockingjay: Unsupervised Speech Representation Learning with Deep Bidirectional Transformer Encoders
- vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations
- Common Voice: A Massively-Multilingual Speech Corpus
- Audio ALBERT: A Lite BERT for Self-supervised Learning of Audio Representation
- Learning to Speak Fluently in a Foreign Language: Multilingual Speech Synthesis and Cross-Language Voice Cloning
- audioLIME: Listenable Explanations Using Source Separation
- Multi-modal Automated Speech Scoring using Attention Fusion
- Do Language Embeddings Capture Scales?
- Towards Modelling Coherence in Spoken Discourse
Cited by in corpus (4)
- Neural Analysis and Synthesis: Reconstructing Speech from Self-Supervised Representations
- Speaker-Conditioned Hierarchical Modeling for Automated Speech Scoring
- AES Systems Are Both Overstable And Oversensitive: Explaining Why And Proposing Defenses
- Probing Acoustic Representations for Phonetic Properties