A Comprehensive Survey of Scene Graphs: Generation and Application
arXiv:2104.01111 · doi:10.1109/TPAMI.2021.3137605
Abstract
Scene graph is a structured representation of a scene that can clearly express the objects, attributes, and relationships between objects in the scene. As computer vision technology continues to develop, people are no longer satisfied with simply detecting and recognizing objects in images; instead, people look forward to a higher level of understanding and reasoning about visual scenes. For example, given an image, we want to not only detect and recognize objects in the image, but also know the relationship between objects (visual relationship detection), and generate a text description (image captioning) based on the image content. Alternatively, we might want the machine to tell us what the little girl in the image is doing (Visual Question Answering (VQA)), or even remove the dog from the image and find similar images (image editing and retrieval), etc. These tasks require a higher level of understanding and reasoning for image vision tasks. The scene graph is just such a powerful tool for scene understanding. Therefore, scene graphs have attracted the attention of a large number of researchers, and related research is often cross-modal, complex, and rapidly developing. However, no relatively systematic survey of scene graphs exists at present. To this end, this survey conducts a comprehensive investigation of the current scene graph research. More specifically, we first summarized the general definition of the scene graph, then conducted a comprehensive and systematic discussion on the generation method of the scene graph (SGG) and the SGG with the aid of prior knowledge. We then investigated the main applications of scene graphs and summarized the most commonly used datasets. Finally, we provide some insights into the future development of scene graphs. We believe this will be a very helpful foundation for future research on scene graphs.
25 pages
References in corpus (14)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Efficient Inference in Fully Connected CRFs with Gaussian Edge Potentials
- On the Properties of Neural Machine Translation: Encoder-Decoder Approaches
- Recurrent Models of Visual Attention
- PCPL: Predicate-Correlation Perception Learning for Unbiased Scene Graph Generation
- Learning to generalize to new compositions in image understanding
- Interactive Image Generation Using Scene Graphs
- Graph Density-Aware Losses for Novel Compositions in Scene Graph Generation
- Natural Language Guided Visual Relationship Detection
- Tackling the Unannotated: Scene Graph Generation with Bias-Reduced Models
- Using Scene Graph Context to Improve Image Generation
- TPsgtR: Neural-Symbolic Tensor Product Scene-Graph-Triplet Representation for Image Captioning
- Scene Graph based Image Retrieval -- A case study on the CLEVR Dataset
- Dual ResGCN for Balanced Scene GraphGeneration
Cited by in corpus (18)
- From SLAM to Situational Awareness: Challenges and Survey
- Efficient Token-Guided Image-Text Retrieval with Consistent Multimodal Contrastive Training
- PlantoGraphy: Incorporating Iterative Design Process into Generative Artificial Intelligence for Landscape Rendering
- RoMFAC: A robust mean-field actor-critic reinforcement learning against adversarial perturbations on states
- Graph Neural Networks in Vision-Language Image Understanding: A Survey
- Artificial Intelligence-based algorithms in medical image scan seg-mentation and intelligent visual-content generation -- a concise overview
- GNNBuilder: An Automated Framework for Generic Graph Neural Network Accelerator Generation, Simulation, and Optimization
- A Universal Knowledge Model and Cognitive Architecture for Prototyping AGI
- Synthesizing Event-centric Knowledge Graphs of Daily Activities Using Virtual Space
- A Reverse Causal Framework to Mitigate Spurious Correlations for Debiasing Scene Graph Generation
- A Survey on Learning from Graphs with Heterophily: Recent Advances and Future Directions
- Zero-Shot Scene Graph Generation via Triplet Calibration and Reduction
- Information-Theoretic Detection of Bimanual Interactions for Dual-Arm Robot Plan Generation
- Combining Optimal Path Search With Task-Dependent Learning in a Neural Network
- HNC: Leveraging Hard Negative Captions towards Models with Fine-Grained Visual-Linguistic Comprehension Capabilities
- From the Laboratory to Real-World Application: Evaluating Zero-Shot Scene Interpretation on Edge Devices for Mobile Robotics
- Implicit Shape Model Trees: Recognition of 3-D Indoor Scenes and Prediction of Object Poses for Mobile Robots
- In-situ Value-aligned Human-Robot Interactions with Physical Constraints