Hollywood in Homes: Crowdsourcing Data Collection for Activity Understanding
arXiv:1604.01753
Abstract
Computer vision has a great potential to help our daily lives by searching for lost keys, watering flowers or reminding us to take a pill. To succeed with such tasks, computer vision methods need to be trained from real and diverse examples of our daily dynamic scenes. While most of such scenes are not particularly exciting, they typically do not appear on YouTube, in movies or TV broadcasts. So how do we collect sufficiently many diverse but boring samples representing our lives? We propose a novel Hollywood in Homes approach to collect such data. Instead of shooting videos in the lab, we ensure diversity by distributing and crowdsourcing the whole process of video creation from script writing to video recording and annotation. Following this procedure we collect a new dataset, Charades, with hundreds of people recording videos in their own homes, acting out casual everyday activities. The dataset is composed of 9,848 annotated videos with an average length of 30 seconds, showing activities of 267 people from three continents. Each video is annotated by multiple free-text descriptions, action labels, action intervals and classes of interacted objects. In total, Charades provides 27,847 video descriptions, 66,500 temporally localized intervals for 157 action classes and 41,104 labels for 46 object classes. Using this rich data, we evaluate and provide baseline results for several tasks including action recognition and automatic description generation. We believe that the realism, diversity, and casual nature of this dataset will present unique challenges and new opportunities for computer vision community.
References in corpus (6)
- Two-Stream Convolutional Networks for Action Recognition in Videos
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- Microsoft COCO Captions: Data Collection and Evaluation Server
- Exploring Nearest Neighbor Approaches for Image Captioning
- Using Descriptive Video Services to Create a Large Data Source for Video Annotation Research
- A Dataset for Movie Description
Cited by in corpus (53)
- Human Action Recognition from Various Data Modalities: A Review
- Attentional Pooling for Action Recognition
- R-C3D: Region Convolutional 3D Network for Temporal Activity Detection
- Charades-Ego: A Large-Scale Dataset of Paired Third and First Person Videos
- CDC: Convolutional-De-Convolutional Networks for Precise Temporal Action Localization in Untrimmed Videos
- The Eighth Dialog System Technology Challenge
- VideoGraph: Recognizing Minutes-Long Human Activities in Videos
- A Better Baseline for AVA
- Learning 2D Temporal Adjacent Networks for Moment Localization with Natural Language
- FedScale: Benchmarking Model and System Performance of Federated Learning at Scale
- Predictive-Corrective Networks for Action Detection
- Shedding Light on Blind Spots: Developing a Reference Architecture to Leverage Video Data for Process Mining
- Parameter Efficient Multimodal Transformers for Video Representation Learning
- Concurrent Activity Recognition with Multimodal CNN-LSTM Structure
- Video Understanding as Machine Translation
- Audio Visual Scene-Aware Dialog (AVSD) Challenge at DSTC7
- Video Captioning via Hierarchical Reinforcement Learning
- LabelSens: Enabling Real-time Sensor Data Labelling at the point of Collection on Edge Computing
- Read, Watch, and Move: Reinforcement Learning for Temporally Grounding Natural Language Descriptions in Videos
- From FiLM to Video: Multi-turn Question Answering with Multi-modal Context
- Multi-step Joint-Modality Attention Network for Scene-Aware Dialogue System
- Video Representation Learning Using Discriminative Pooling
- Weakly-Supervised Video Moment Retrieval via Semantic Completion Network
- VirtualHome: Simulating Household Activities via Programs
- Temporal Dynamic Graph LSTM for Action-driven Video Object Detection
- Attend and Interact: Higher-Order Object Interactions for Video Understanding
- IMUTube: Automatic Extraction of Virtual on-body Accelerometry from Video for Human Activity Recognition
- Asynchronous Temporal Fields for Action Recognition
- Context, Attention and Audio Feature Explorations for Audio Visual Scene-Aware Dialog
- Relational Action Forecasting
- PIC: Permutation Invariant Convolution for Recognizing Long-range Activities
- Joint Discovery of Object States and Manipulation Actions
- A Survey on Natural Language Video Localization
- Temporal Query Networks for Fine-grained Video Understanding
- VidTr: Video Transformer Without Convolutions
- Trespassing the Boundaries: Labeling Temporal Bounds for Object Interactions in Egocentric Video
- ActionVLAD: Learning spatio-temporal aggregation for action classification
- Tree-Structured Policy based Progressive Reinforcement Learning for Temporally Language Grounding in Video
- Towards an Unequivocal Representation of Actions
- TimeGate: Conditional Gating of Segments in Long-range Activities
- Video Moment Retrieval with Text Query Considering Many-to-Many Correspondence Using Potentially Relevant Pair
- Team RUC_AIM3 Technical Report at ActivityNet 2021: Entities Object Localization
- PyTorchVideo: A Deep Learning Library for Video Understanding
- Improving Classification by Improving Labelling: Introducing Probabilistic Multi-Label Object Interaction Recognition
- Deconfounded Video Moment Retrieval with Causal Intervention
- Follow the Attention: Combining Partial Pose and Object Motion for Fine-Grained Action Detection
- Anticipating Daily Intention using On-Wrist Motion Triggered Sensing
- TAN: Temporal Aggregation Network for Dense Multi-label Action Recognition
- Recurrent Residual Module for Fast Inference in Videos
- Delta Sampling R-BERT for limited data and low-light action recognition
- A Continuous, Full-scope, Spatio-temporal Tracking Metric based on KL-divergence
- Attentive Sequence to Sequence Translation for Localizing Clips of Interest by Natural Language Descriptions
- A Mobile Robot Generating Video Summaries of Seniors' Indoor Activities