How2: A Large-scale Dataset for Multimodal Language Understanding
arXiv:1811.00347
Abstract
In this paper, we introduce How2, a multimodal collection of instructional videos with English subtitles and crowdsourced Portuguese translations. We also present integrated sequence-to-sequence baselines for machine translation, automatic speech recognition, spoken language translation, and multimodal summarization. By making available data and code for several multimodal natural language tasks, we hope to stimulate more research on these and similar challenges, to obtain a deeper understanding of multimodality in language processing.
References in corpus (5)
- Adam: A Method for Stochastic Optimization
- Teaching Machines to Read and Comprehend
- Augmenting Librispeech with French Translations: A Multimodal Corpus for Direct Speech Translation Evaluation
- Sequence-to-Sequence Models Can Directly Translate Foreign Speech
- Visual Features for Context-Aware Speech Recognition
Cited by in corpus (26)
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips
- Abstractive Text Summarization: State of the Art, Challenges, and Improvements
- Trends in Integration of Vision and Language Research: A Survey of Tasks, Datasets, and Methods
- VATEX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language Research
- Abstractive Summarization of Spoken and Written Instructions with BERT
- Beyond Voice Activity Detection: Hybrid Audio Segmentation for Direct Speech Translation
- Low-Latency Sequence-to-Sequence Speech Recognition and Translation by Partial Hypothesis Selection
- ESPnet-ST: All-in-One Speech Translation Toolkit
- Predicting Actions to Help Predict Translations
- CUPID: Adaptive Curation of Pre-training Data for Video-and-Language Representation Learning
- Grounding Object Detections With Transcriptions
- Analyzing Utility of Visual Context in Multimodal Speech Recognition Under Noisy Conditions
- Strategies for improving low resource speech to text translation relying on pre-trained ASR models
- Multimodal Speech Recognition for Language-Guided Embodied Agents
- Gender in Danger? Evaluating Speech Translation Technology on the MuST-SHE Corpus
- A Benchmark for Structured Procedural Knowledge Extraction from Cooking Videos
- Vision Guided Generative Pre-trained Language Models for Multimodal Abstractive Summarization
- TMT: A Transformer-based Modal Translator for Improving Multimodal Sequence Representations in Audio Visual Scene-aware Dialog
- Rudder: A Cross Lingual Video and Text Retrieval Dataset
- Routing with Self-Attention for Multimodal Capsule Networks
- Adaptive Beam Search to Enhance On-device Abstractive Summarization
- Attention-based Multi-hypothesis Fusion for Speech Summarization
- Speech Summarization using Restricted Self-Attention
- TCT: A Cross-supervised Learning Method for Multimodal Sequence Representation
- ON-TRAC Consortium End-to-End Speech Translation Systems for the IWSLT 2019 Shared Task
- Impact of Encoding and Segmentation Strategies on End-to-End Simultaneous Speech Translation