Analyzing analytical methods: The case of phonology in neural models of spoken language
arXiv:2004.07070 · doi:10.18653/v1/2020.acl-main.381
Abstract
Given the fast development of analysis techniques for NLP and speech processing systems, few systematic studies have been conducted to compare the strengths and weaknesses of each method. As a step in this direction we study the case of representations of phonology in neural network models of spoken language. We use two commonly applied analytical techniques, diagnostic classifiers and representational similarity analysis, to quantify to what extent neural activation patterns encode phonemes and phoneme sequences. We manipulate two factors that can affect the outcome of analysis. First, we investigate the role of learning by comparing neural activations extracted from trained versus randomly-initialized models. Second, we examine the temporal scope of the activations by probing both local activations corresponding to a few milliseconds of the speech signal, and global activations pooled over the whole utterance. We conclude that reporting analysis results with randomly initialized models is crucial, and that global-scope methods tend to yield more consistent results and we recommend their use as a complement to local-scope diagnostic methods.
ACL 2020
References in corpus (7)
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- ADADELTA: An Adaptive Learning Rate Method
- SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition
- Fine-grained Analysis of Sentence Embeddings Using Auxiliary Prediction Tasks
- Correlating neural and symbolic representations of language
- Language learning using Speech to Image retrieval
- Analyzing Hidden Representations in End-to-End Automatic Speech Recognition Systems
Cited by in corpus (4)
- Visually grounded models of spoken language: A survey of datasets, architectures and evaluation techniques
- Wave to Syntax: Probing spoken language models for syntax
- Human-like Linguistic Biases in Neural Speech Models: Phonetic Categorization and Phonotactic Constraints in Wav2Vec2.0
- Discrete representations in neural models of spoken language