Speech2Vec: A Sequence-to-Sequence Framework for Learning Word Embeddings from Speech
arXiv:1803.08976
Abstract
In this paper, we propose a novel deep neural network architecture, Speech2Vec, for learning fixed-length vector representations of audio segments excised from a speech corpus, where the vectors contain semantic information pertaining to the underlying spoken words, and are close to other vectors in the embedding space if their corresponding underlying spoken words are semantically similar. The proposed model can be viewed as a speech version of Word2Vec. Its design is based on a RNN Encoder-Decoder framework, and borrows the methodology of skipgrams or continuous bag-of-words for training. Learning word embeddings directly from speech enables Speech2Vec to make use of the semantic information carried by speech that does not exist in plain text. The learned word embeddings are evaluated and analyzed on 13 widely used word similarity benchmarks, and outperform word embeddings learned by Word2Vec from the transcriptions.
Accepted to Interspeech 2018; camera-ready version. arXiv admin note: text overlap with arXiv:1711.01515
References in corpus (7)
- Sequence to Sequence Learning with Neural Networks
- Learning General Purpose Distributed Sentence Representations via Large Scale Multi-task Learning
- A segmental framework for fully-unsupervised large-vocabulary speech recognition
- Discriminative Acoustic Word Embeddings: Recurrent Neural Network-Based Approaches
- Learning Word Embeddings from Speech
- Speech-Based Visual Question Answering
- Character Composition Model with Convolutional Neural Networks for Dependency Parsing on Morphologically Rich Languages
Cited by in corpus (15)
- Libri-Light: A Benchmark for ASR with Limited or No Supervision
- Pushing the Limits of Semi-Supervised Learning for Automatic Speech Recognition
- Deep Learning Techniques for Future Intelligent Cross-Media Retrieval
- BigSSL: Exploring the Frontier of Large-Scale Semi-Supervised Learning for Automatic Speech Recognition
- Unsupervised Cross-Modal Alignment of Speech and Text Embedding Spaces
- Learning Household Task Knowledge from WikiHow Descriptions
- Supervised and Unsupervised Transfer Learning for Question Answering
- Neuralogram: A Deep Neural Network Based Representation for Audio Signals
- L2RS: A Learning-to-Rescore Mechanism for Automatic Speech Recognition
- Neural Methods for Point-wise Dependency Estimation
- DP-Parse: Finding Word Boundaries from Raw Speech with an Instance Lexicon
- Supervised Speaker Embedding De-Mixing in Two-Speaker Environment
- Telephonetic: Making Neural Language Models Robust to ASR and Semantic Noise
- How Does That Sound? Multi-Language SpokenName2Vec Algorithm Using Speech Generation and Deep Learning
- Cetacean Translation Initiative: a roadmap to deciphering the communication of sperm whales