SPEECH-COCO: 600k Visually Grounded Spoken Captions Aligned to MSCOCO Data Set
arXiv:1707.08435 · doi:10.21437/GLU.2017-9
Abstract
This paper presents an augmentation of MSCOCO dataset where speech is added to image and text. Speech captions are generated using text-to-speech (TTS) synthesis resulting in 616,767 spoken captions (more than 600h) paired with images. Disfluencies and speed perturbation are added to the signal in order to sound more natural. Each speech signal (WAV) is paired with a JSON file containing exact timecode for each word/syllable/phoneme in the spoken caption. Such a corpus could be used for Language and Vision (LaVi) tasks including speech input or output instead of text. Investigating multimodal learning schemes for unsupervised speech pattern discovery is also possible with this corpus, as demonstrated by a preliminary study conducted on a subset of the corpus (10h, 10k spoken captions). The dataset is available on Zenodo: https://zenodo.org/record/4282267
Data set available on https://zenodo.org/record/4282267. Presented at GLU (Grounded Language Understanding) Satellite Workshop of Interspeech 2017
References in corpus (5)
- Attention-Based Models for Speech Recognition
- Microsoft COCO Captions: Data Collection and Evaluation Server
- Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations
- Representations of language in a model of visually grounded speech signal
- Visual Question Answering: A Survey of Methods and Datasets
Cited by in corpus (5)
- Visually grounded models of spoken language: A survey of datasets, architectures and evaluation techniques
- Evaluation of Audio-Visual Alignments in Visually Grounded Speech Models
- Models of Visually Grounded Speech Signal Pay Attention To Nouns: a Bilingual Experiment on English and Japanese
- Spoken ObjectNet: A Bias-Controlled Spoken Caption Dataset
- Empowering cyberphysical systems of systems with intelligence