A segmental framework for fully-unsupervised large-vocabulary speech recognition
arXiv:1606.06950 · doi:10.1016/j.csl.2017.04.008
Abstract
Zero-resource speech technology is a growing research area that aims to develop methods for speech processing in the absence of transcriptions, lexicons, or language modelling text. Early term discovery systems focused on identifying isolated recurring patterns in a corpus, while more recent full-coverage systems attempt to completely segment and cluster the audio into word-like units---effectively performing unsupervised speech recognition. This article presents the first attempt we are aware of to apply such a system to large-vocabulary multi-speaker data. Our system uses a Bayesian modelling framework with segmental word representations: each word segment is represented as a fixed-dimensional acoustic embedding obtained by mapping the sequence of feature frames to a single embedding vector. We compare our system on English and Xitsonga datasets to state-of-the-art baselines, using a variety of measures including word error rate (obtained by mapping the unsupervised output to ground truth transcriptions). Very high word error rates are reported---in the order of 70--80% for speaker-dependent and 80--95% for speaker-independent systems---highlighting the difficulty of this task. Nevertheless, in terms of cluster quality and word segmentation metrics, we show that by imposing a consistent top-down segmentation while also using bottom-up knowledge from detected syllable boundaries, both single-speaker and multi-speaker versions of our system outperform a purely bottom-up single-speaker syllable-based approach. We also show that the discovered clusters can be made less speaker- and gender-specific by using an unsupervised autoencoder-like feature extractor to learn better frame-level features (prior to embedding). Our system's discovered clusters are still less pure than those of unsupervised term discovery systems, but provide far greater coverage.
15 pages, 6 figures, 8 tables
References in corpus (1)
Cited by in corpus (26)
- Unsupervised speech representation learning using WaveNet autoencoders
- Unsupervised Cross-Modal Alignment of Speech and Text Embedding Spaces
- Unsupervised Automatic Speech Recognition: A Review
- CiwGAN and fiwGAN: Encoding information in acoustic data to model lexical learning with Generative Adversarial Networks
- Speech2Vec: A Sequence-to-Sequence Framework for Learning Word Embeddings from Speech
- Self-supervised language learning from raw audio: Lessons from the Zero Resource Speech Challenge
- Word Segmentation on Discovered Phone Units with Dynamic Programming and Self-Supervised Scoring
- Unsupervised Speech Recognition
- Learning Word Embeddings from Speech
- Unsupervised feature learning for speech using correspondence and Siamese networks
- Learning Hierarchical Discrete Linguistic Units from Visually-Grounded Speech
- Multilingual and Unsupervised Subword Modeling for Zero-Resource Languages
- Exploring TTS without T Using Biologically/Psychologically Motivated Neural Network Modules (ZeroSpeech 2020)
- Double Articulation Analyzer with Prosody for Unsupervised Word and Phoneme Discovery
- Unsupervised neural and Bayesian models for zero-resource speech processing
- Segmental Audio Word2Vec: Representing Utterances as Sequences of Vectors with Applications in Spoken Term Detection
- Unsupervised Subword Modeling Using Autoregressive Pretraining and Cross-Lingual Phone-Aware Modeling
- Sequence Prediction with Neural Segmental Models
- Completely Unsupervised Phoneme Recognition by Adversarially Learning Mapping Relationships from Audio Embeddings
- Unsupervised Pattern Discovery from Thematic Speech Archives Based on Multilingual Bottleneck Features
- Learning Joint Acoustic-Phonetic Word Embeddings
- From Semi-supervised to Almost-unsupervised Speech Recognition with Very-low Resource by Jointly Learning Phonetic Structures from Audio and Text Embeddings
- Towards Visually Grounded Sub-Word Speech Unit Discovery
- Bayesian Subspace HMM for the Zerospeech 2020 Challenge
- Unsupervised Spoken Term Discovery on Untranscribed Speech
- Phonetic-and-Semantic Embedding of Spoken Words with Applications in Spoken Content Retrieval