Representations of language in a model of visually grounded speech signal
arXiv:1702.01991 · doi:10.18653/v1/P17-1057
Abstract
We present a visually grounded model of speech perception which projects spoken utterances and images to a joint semantic space. We use a multi-layer recurrent highway network to model the temporal nature of spoken speech, and show that it learns to extract both form and meaning-based linguistic knowledge from the input signal. We carry out an in-depth analysis of the representations used by different components of the trained model and show that encoding of semantic aspects tends to become richer as we go up the hierarchy of layers, whereas encoding of form-related aspects of the language input tends to initially increase and then plateau or decrease.
Accepted at ACL 2017
References in corpus (9)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Pixel Recurrent Neural Networks
- Understanding Neural Networks through Representation Erasure
- Fine-grained Analysis of Sentence Embeddings Using Auxiliary Prediction Tasks
- Order-Embeddings of Images and Language
- From phonemes to images: levels of representation in a recurrent neural model of visually-grounded language learning
- Automatic Description Generation from Images: A Survey of Models, Datasets, and Evaluation Measures
- Learning language through pictures
- Memory Visualization for Gated Recurrent Neural Networks in Speech Recognition
Cited by in corpus (36)
- Self-Supervised Speech Representation Learning: A Review
- Effectiveness of self-supervised pre-training for speech recognition
- Unsupervised Automatic Speech Recognition: A Review
- Language learning using Speech to Image retrieval
- Trends in Integration of Vision and Language Research: A Survey of Tasks, Datasets, and Methods
- Visually grounded models of spoken language: A survey of datasets, architectures and evaluation techniques
- Analyzing Hidden Representations in End-to-End Automatic Speech Recognition Systems
- Self-supervised language learning from raw audio: Lessons from the Zero Resource Speech Challenge
- SPEECH-COCO: 600k Visually Grounded Spoken Captions Aligned to MSCOCO Data Set
- Analyzing analytical methods: The case of phonology in neural models of spoken language
- Voice-assisted Image Labelling for Endoscopic Ultrasound Classification using Neural Networks
- Speech-Image Semantic Alignment Does Not Depend on Any Prior Classification Tasks
- improving partition-block-based acoustic echo canceler in under-modeling scenarios
- Learning English with Peppa Pig
- Learning semantic sentence representations from visually grounded language without lexical knowledge
- Learning Hierarchical Discrete Linguistic Units from Visually-Grounded Speech
- Revisiting the Hierarchical Multiscale LSTM
- ZR-2021VG: Zero-Resource Speech Challenge, Visually-Grounded Language Modelling track, 2021 edition
- Double Articulation Analyzer with Prosody for Unsupervised Word and Phoneme Discovery
- Keyword localisation in untranscribed speech using visually grounded speech models
- Symbolic inductive bias for visually grounded learning of spoken language
- Talk, Don't Write: A Study of Direct Speech-Based Image Retrieval
- Evaluation of Audio-Visual Alignments in Visually Grounded Speech Models
- Acoustic Feature Learning via Deep Variational Canonical Correlation Analysis
- Semantic speech retrieval with a visually grounded model of untranscribed speech
- Revisiting Cross Modal Retrieval
- Towards localisation of keywords in speech using weak supervision
- Learning to Discover, Ground and Use Words with Segmental Neural Language Models
- Unsupervised Multimodal Word Discovery based on Double Articulation Analysis with Co-occurrence cues
- Semantic query-by-example speech search using visual grounding
- Simultaneous or Sequential Training? How Speech Representations Cooperate in a Multi-Task Self-Supervised Learning System
- Grounding 'Grounding' in NLP
- Cascaded Multilingual Audio-Visual Learning from Videos
- Spoken ObjectNet: A Bias-Controlled Spoken Caption Dataset
- Interpretable Textual Neuron Representations for NLP
- On the difficulty of a distributional semantics of spoken language