Contextual Action Recognition with R*CNN
arXiv:1505.01197
Abstract
There are multiple cues in an image which reveal what action a person is performing. For example, a jogger has a pose that is characteristic for jogging, but the scene (e.g. road, trail) and the presence of other joggers can be an additional source of information. In this work, we exploit the simple observation that actions are accompanied by contextual cues to build a strong action recognition system. We adapt RCNN to use more than one region for classification while still maintaining the ability to localize the action. We call our system R*CNN. The action-specific models and the feature maps are trained jointly, allowing for action specific representations to emerge. R*CNN achieves 90.2% mean AP on the PASAL VOC Action dataset, outperforming all other approaches in the field by a significant margin. Last, we show that R*CNN is not limited to action recognition. In particular, R*CNN can also be used to tackle fine-grained tasks such as attribute classification. We validate this claim by reporting state-of-the-art performance on the Berkeley Attributes of People dataset.
References in corpus (4)
Cited by in corpus (37)
- Improving Person Re-identification by Attribute and Identity Learning
- Attentional Pooling for Action Recognition
- Temporal Activity Detection in Untrimmed Videos with Recurrent Neural Networks
- Detecting Visual Relationships with Deep Relational Networks
- Spatial Transformer Networks
- Detecting and Recognizing Human-Object Interactions
- Learning Spatial Regularization with Image-level Supervisions for Multi-label Image Classification
- A Survey on Deep Learning Methods for Robot Vision
- Be Your Own Prada: Fashion Synthesis with Structural Coherence
- HAKE: Human Activity Knowledge Engine
- Attribute Recognition by Joint Recurrent Learning of Context and Correlation
- Context-Aware Self-Attention Networks
- Exploring Person Context and Local Scene Context for Object Detection
- LabelBank: Revisiting Global Perspectives for Semantic Segmentation
- Generating Descriptions with Grounded and Co-Referenced People
- Spatio-Temporal Action Detection with Cascade Proposal and Location Anticipation
- Weakly Supervised Representation Learning for Unsynchronized Audio-Visual Events
- Attend and Interact: Higher-Order Object Interactions for Video Understanding
- Im2Flow: Motion Hallucination from Static Images for Action Recognition
- The Role of Typicality in Object Classification: Improving The Generalization Capacity of Convolutional Neural Networks
- Spatio-temporal Human Action Localisation and Instance Segmentation in Temporally Untrimmed Videos
- Dual-Glance Model for Deciphering Social Relationships
- Is Object Detection Necessary for Human-Object Interaction Recognition?
- Agent-Centric Risk Assessment: Accident Anticipation and Risky Region Localization
- ORD: Object Relationship Discovery for Visual Dialogue Generation
- Low-Latency Human Action Recognition with Weighted Multi-Region Convolutional Neural Network
- Spatial Memory for Context Reasoning in Object Detection
- Improved Hard Example Mining by Discovering Attribute-based Hard Person Identity
- Deep Imbalanced Attribute Classification using Visual Attention Aggregation
- Learning to Recognize Objects by Retaining other Factors of Variation
- Story-oriented Image Selection and Placement
- Understanding and Predicting The Attractiveness of Human Action Shot
- Learning Action Concept Trees and Semantic Alignment Networks from Image-Description Data
- Proposal Flow: Semantic Correspondences from Object Proposals
- Deep neural networks can be improved using human-derived contextual expectations
- Single Image Action Recognition using Semantic Body Part Actions
- Loss Guided Activation for Action Recognition in Still Images