Visually grounded models of spoken language: A survey of datasets, architectures and evaluation techniques
arXiv:2104.13225 · doi:10.1613/jair.1.12967
Abstract
This survey provides an overview of the evolution of visually grounded models of spoken language over the last 20 years. Such models are inspired by the observation that when children pick up a language, they rely on a wide range of indirect and noisy clues, crucially including signals from the visual modality co-occurring with spoken utterances. Several fields have made important contributions to this approach to modeling or mimicking the process of learning language: Machine Learning, Natural Language and Speech Processing, Computer Vision and Cognitive Science. The current paper brings together these contributions in order to provide a useful introduction and overview for practitioners in all these areas. We discuss the central research questions addressed, the timeline of developments, and the datasets which enabled much of this work. We then summarize the main modeling architectures and offer an exhaustive overview of the evaluation metrics and analysis techniques.
References in corpus (13)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
- YouTube-8M: A Large-Scale Video Classification Benchmark
- Deep Fragment Embeddings for Bidirectional Image Sentence Mapping
- Analyzing Hidden Representations in End-to-End Automatic Speech Recognition Systems
- From phonemes to images: levels of representation in a recurrent neural model of visually-grounded language learning
- Room-Across-Room: Multilingual Vision-and-Language Navigation with Dense Spatiotemporal Grounding
- Learning Hierarchical Discrete Linguistic Units from Visually-Grounded Speech
- ZR-2021VG: Zero-Resource Speech Challenge, Visually-Grounded Language Modelling track, 2021 edition
- Can phones, syllables, and words emerge as side-products of cross-situational audiovisual learning? -- A computational investigation
- Fast-Slow Transformer for Visually Grounding Speech
- Talk, Don't Write: A Study of Direct Speech-Based Image Retrieval
- Towards localisation of keywords in speech using weak supervision
Cited by in corpus (10)
- Self-Supervised Speech Representation Learning: A Review
- ConceptBeam: Concept Driven Target Speech Extraction
- Learning English with Peppa Pig
- ZR-2021VG: Zero-Resource Speech Challenge, Visually-Grounded Language Modelling track, 2021 edition
- Keyword localisation in untranscribed speech using visually grounded speech models
- Video-Guided Curriculum Learning for Spoken Video Grounding
- Wave to Syntax: Probing spoken language models for syntax
- Simultaneous or Sequential Training? How Speech Representations Cooperate in a Multi-Task Self-Supervised Learning System
- Discrete representations in neural models of spoken language
- Cascaded Multilingual Audio-Visual Learning from Videos