Towards Automatic Learning of Procedures from Web Instructional Videos
arXiv:1703.09788
Abstract
The potential for agents, whether embodied or software, to learn by observing other agents performing procedures involving objects and actions is rich. Current research on automatic procedure learning heavily relies on action labels or video subtitles, even during the evaluation phase, which makes them infeasible in real-world scenarios. This leads to our question: can the human-consensus structure of a procedure be learned from a large set of long, unconstrained videos (e.g., instructional videos from YouTube) with only visual evidence? To answer this question, we introduce the problem of procedure segmentation--to segment a video procedure into category-independent procedure segments. Given that no large-scale dataset is available for this problem, we collect a large-scale procedure segmentation dataset with procedure segments temporally localized and described; we use cooking videos and name the dataset YouCook2. We propose a segment-level recurrent network for generating procedure segments by modeling the dependencies across segments. The generated segments can be used as pre-processing for other tasks, such as dense video captioning and event parsing. We show in our experiments that the proposed model outperforms competitive baselines in procedure segmentation.
AAAI 2018 Camera-ready version. See http://youcook2.eecs.umich.edu for YouCook2 dataset
Cited by in corpus (79)
- Transformers in Vision: A Survey
- HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips
- VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text
- Learning Video Representations using Contrastive Bidirectional Transformer
- Multimodal Co-learning: Challenges, Applications with Datasets, Recent Advances and Future Directions
- Self-Supervised MultiModal Versatile Networks
- Unifying Vision-and-Language Tasks via Text Generation
- VideoBERT: A Joint Model for Video and Language Representation Learning
- Self-Monitoring Navigation Agent via Auxiliary Progress Estimation
- VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling
- Generative Adversarial Imitation from Observation
- Weakly-Supervised Video Object Grounding from Text by Loss Weighting and Object Interaction
- COOT: Cooperative Hierarchical Transformer for Video-Text Representation Learning
- TricorNet: A Hybrid Temporal Convolutional and Recurrent Network for Video Action Segmentation
- HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training
- Trends in Integration of Vision and Language Research: A Survey of Tasks, Datasets, and Methods
- Less is More: ClipBERT for Video-and-Language Learning via Sparse Sampling
- VideoGraph: Recognizing Minutes-Long Human Activities in Videos
- Comprehensive Instructional Video Analysis: The COIN Dataset and Performance Evaluation
- End-to-End Dense Video Captioning with Masked Transformer
- Visually grounded models of spoken language: A survey of datasets, architectures and evaluation techniques
- PlaTe: Visually-Grounded Planning with Transformers in Procedural Tasks
- VATEX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language Research
- VALUE: A Multi-Task Benchmark for Video-and-Language Understanding Evaluation
- Support-set bottlenecks for video-text representation learning
- Drop-DTW: Aligning Common Signal Between Sequences While Dropping Outliers
- Neural Language Modeling with Visual Features
- Deep Learning for Vision-based Prediction: A Survey
- TVR: A Large-Scale Dataset for Video-Subtitle Moment Retrieval
- Learning Video Representations from Textual Web Supervision
- Multimodal Pretraining for Dense Video Captioning
- COIN: A Large-scale Dataset for Comprehensive Instructional Video Analysis
- Look at What I'm Doing: Self-Supervised Spatial Grounding of Narrations in Instructional Videos
- MART: Memory-Augmented Recurrent Transformer for Coherent Video Paragraph Captioning
- Mining YouTube - A dataset for learning fine-grained action concepts from webly supervised video data
- Attention is all you need for Videos: Self-attention based Video Summarization using Universal Transformers
- Procedure Planning in Instructional Videos
- The End-of-End-to-End: A Video Understanding Pentathlon Challenge (2020)
- ZR-2021VG: Zero-Resource Speech Challenge, Visually-Grounded Language Modelling track, 2021 edition
- VX2TEXT: End-to-End Learning of Video-Based Text Generation From Multimodal Inputs
- A Better Use of Audio-Visual Cues: Dense Video Captioning with Bi-modal Transformer
- Induce, Edit, Retrieve: Language Grounded Multimodal Schema for Instructional Video Retrieval
- Video-Text Pre-training with Learned Regions
- CUPID: Adaptive Curation of Pre-training Data for Video-and-Language Representation Learning
- Grounding Object Detections With Transcriptions
- VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
- VLM: Task-agnostic Video-Language Model Pre-training for Video Understanding
- Advancing High-Resolution Video-Language Representation with Large-Scale Video Transcriptions
- Enabling Robots to Understand Incomplete Natural Language Instructions Using Commonsense Reasoning
- A Benchmark for Structured Procedural Knowledge Extraction from Cooking Videos
- TAB-VCR: Tags and Attributes based Visual Commonsense Reasoning Baselines
- TACo: Token-aware Cascade Contrastive Learning for Video-Text Alignment
- SACT: Self-Aware Multi-Space Feature Composition Transformer for Multinomial Attention for Video Captioning
- Winning the ICCV'2021 VALUE Challenge: Task-aware Ensemble and Transfer Learning with Visual Concepts
- CLIP4Caption ++: Multi-CLIP for Video Caption
- Structure-Aware Generation Network for Recipe Generation from Images
- Learning to Segment Actions from Observation and Narration
- ViSeRet: A simple yet effective approach to moment retrieval via fine-grained video segmentation
- DVCFlow: Modeling Information Flow Towards Human-like Video Captioning
- Survey: Transformer based Video-Language Pre-training
- Learning Temporal Dynamics from Cycles in Narrated Video
- Video Caption Dataset for Describing Human Actions in Japanese
- Zero-Shot Anticipation for Instructional Activities
- ActBERT: Learning Global-Local Video-Text Representations
- Designing Multimodal Datasets for NLP Challenges
- Reparameterized Variational Divergence Minimization for Stable Imitation
- Routing with Self-Attention for Multimodal Capsule Networks
- Goal-driven text descriptions for images
- Overview of Tencent Multi-modal Ads Video Understanding Challenge
- EVOQUER: Enhancing Temporal Grounding with Video-Pivoted BackQuery Generation
- Neuro-Symbolic Representations for Video Captioning: A Case for Leveraging Inductive Biases for Vision and Language
- GEM: A General Evaluation Benchmark for Multimodal Tasks
- Reconstructing and grounding narrated instructional videos in 3D
- Cascaded Multilingual Audio-Visual Learning from Videos
- Video-aided Unsupervised Grammar Induction
- CrossCLR: Cross-modal Contrastive Learning For Multi-modal Video Representations
- Understanding in Artificial Intelligence
- SwAMP: Swapped Assignment of Multi-Modal Pairs for Cross-Modal Retrieval
- Video2Skill: Adapting Events in Demonstration Videos to Skills in an Environment using Cyclic MDP Homomorphisms