Pixels to Graphs by Associative Embedding
arXiv:1706.07365
Abstract
Graphs are a useful abstraction of image content. Not only can graphs represent details about individual objects in a scene but they can capture the interactions between pairs of objects. We present a method for training a convolutional neural network such that it takes in an input image and produces a full graph definition. This is done end-to-end in a single stage with the use of associative embeddings. The network learns to simultaneously identify all of the elements that make up a graph and piece them together. We benchmark on the Visual Genome dataset, and demonstrate state-of-the-art performance on the challenging task of scene graph generation.
Updated numbers. Code and pretrained models available at https://github.com/umich-vl/px2graph
References in corpus (10)
- Visual Relationship Detection with Language Priors
- Scene Graph Generation by Iterative Message Passing
- Discovering objects and their relations from entangled scene representations
- Detecting Visual Relationships with Deep Relational Networks
- Visual Translation Embedding Network for Visual Relation Detection
- Learning to generalize to new compositions in image understanding
- Deep Variation-structured Reinforcement Learning for Visual Relationship and Attribute Detection
- Learning to Detect Human-Object Interactions
- Modeling Relationships in Referential Expressions with Compositional Modular Networks
- ViP-CNN: Visual Phrase Guided Convolutional Neural Network
Cited by in corpus (32)
- A Comprehensive Survey of Scene Graphs: Generation and Application
- CornerNet: Detecting Objects as Paired Keypoints
- 3D-BEVIS: Bird's-Eye-View Instance Segmentation
- PasteGAN: A Semi-Parametric Method to Generate Image from Scene Graph
- An Empirical Study on Leveraging Scene Graphs for Visual Question Answering
- Learning to Compose Dynamic Tree Structures for Visual Contexts
- Learning Visual Relation Priors for Image-Text Matching and Image Captioning with Neural Scene Graph Generators
- PPDM: Parallel Point Detection and Matching for Real-time Human-Object Interaction Detection
- Visual Relationship Detection with Visual-Linguistic Knowledge from Multimodal Representations
- Counterfactual Critic Multi-Agent Training for Scene Graph Generation
- Bridging Knowledge Graphs to Generate Scene Graphs
- Large-Scale Visual Relationship Understanding
- Unsupervised Traffic Scene Generation with Synthetic 3D Scene Graphs
- The Limited Multi-Label Projection Layer
- Attentive Relational Networks for Mapping Images to Scene Graphs
- A Multi-Stage Multi-Task Neural Network for Aerial Scene Interpretation and Geolocalization
- VrR-VG: Refocusing Visually-Relevant Relationships
- Graph R-CNN for Scene Graph Generation
- GPS-Net: Graph Property Sensing Network for Scene Graph Generation
- An Interpretable Model for Scene Graph Generation
- Radar Camera Fusion via Representation Learning in Autonomous Driving
- LinkNet: Relational Embedding for Scene Graph
- Relation Transformer Network
- Compositional Video Synthesis with Action Graphs
- StarMap for Category-Agnostic Keypoint and Viewpoint Estimation
- Learning Predicates as Functions to Enable Few-shot Scene Graph Prediction
- Learning latent causal graphs via mixture oracles
- Learning Segmentation Masks with the Independence Prior
- Image-Level Attentional Context Modeling Using Nested-Graph Neural Networks
- Weakly Supervised Visual Semantic Parsing
- Target-Tailored Source-Transformation for Scene Graph Generation
- Generating Videos of Zero-Shot Compositions of Actions and Objects