Learning to Detect Human-Object Interactions
arXiv:1702.05448
Abstract
We study the problem of detecting human-object interactions (HOI) in static images, defined as predicting a human and an object bounding box with an interaction class label that connects them. HOI detection is a fundamental problem in computer vision as it provides semantic information about the interactions among the detected objects. We introduce HICO-DET, a new large benchmark for HOI detection, by augmenting the current HICO classification benchmark with instance annotations. To solve the task, we propose Human-Object Region-based Convolutional Neural Networks (HO-RCNN). At the core of our HO-RCNN is the Interaction Pattern, a novel DNN input that characterizes the spatial relations between two bounding boxes. Experiments on HICO-DET demonstrate that our HO-RCNN, by exploiting human-object spatial relations through Interaction Patterns, significantly improves the performance of HOI detection over baseline approaches.
Accepted in WACV 2018
References in corpus (7)
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Going Deeper with Convolutions
- Visual Semantic Role Labeling
- Visual Relationship Detection with Language Priors
- Detecting Visual Relationships with Deep Relational Networks
- Visual Translation Embedding Network for Visual Relation Detection
- Deep Variation-structured Reinforcement Learning for Visual Relationship and Attribute Detection
Cited by in corpus (15)
- A Comprehensive Survey of Scene Graphs: Generation and Application
- Pixels to Graphs by Associative Embedding
- Visual Entailment: A Novel Task for Fine-Grained Image Understanding
- Detecting and Recognizing Human-Object Interactions
- Knowledge-Embedded Routing Network for Scene Graph Generation
- Attend and Interact: Higher-Order Object Interactions for Video Understanding
- Understanding Human Hands in Contact at Internet Scale
- Relational Action Forecasting
- Glance and Gaze: Inferring Action-aware Points for One-Stage Human-Object Interaction Detection
- LinkNet: Relational Embedding for Scene Graph
- Grounded Objects and Interactions for Video Captioning
- Turbo Learning Framework for Human-Object Interactions Recognition and Human Pose Estimation
- Spatial Priming for Detecting Human-Object Interactions
- Learning Actor Relation Graphs for Group Activity Recognition
- A Tracking System For Baseball Game Reconstruction