Unbiased Scene Graph Generation from Biased Training
arXiv:2002.11949
Abstract
Today's scene graph generation (SGG) task is still far from practical, mainly due to the severe training bias, e.g., collapsing diverse "human walk on / sit on / lay on beach" into "human on beach". Given such SGG, the down-stream tasks such as VQA can hardly infer better scene structures than merely a bag of objects. However, debiasing in SGG is not trivial because traditional debiasing methods cannot distinguish between the good and bad bias, e.g., good context prior (e.g., "person read book" rather than "eat") and bad long-tailed bias (e.g., "near" dominating "behind / in front of"). In this paper, we present a novel SGG framework based on causal inference but not the conventional likelihood. We first build a causal graph for SGG, and perform traditional biased training with the graph. Then, we propose to draw the counterfactual causality from the trained graph to infer the effect from the bad bias, which should be removed. In particular, we use Total Direct Effect (TDE) as the proposed final predicate score for unbiased SGG. Note that our framework is agnostic to any SGG model and thus can be widely applied in the community who seeks unbiased predictions. By using the proposed Scene Graph Diagnosis toolkit on the SGG benchmark Visual Genome and several prevailing models, we observed significant improvements over the previous state-of-the-art methods.
This paper is accepted by CVPR 2020. The code is publicly available on GitHub: https://github.com/KaihuaTang/Scene-Graph-Benchmark.pytorch
References in corpus (9)
- Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
- Counterfactual Fairness
- ImageNet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness
- RUBi: Reducing Unimodal Biases in Visual Question Answering
- Women also Snowboard: Overcoming Bias in Captioning Models
- Influence of Resampling on Accuracy of Imbalanced Classification
- Causal Induction from Visual Observations for Goal Directed Tasks
- Counterfactual VQA: A Cause-Effect Look at Language Bias
- Two Causal Principles for Improving Visual Dialog
Cited by in corpus (14)
- A Comprehensive Survey of Scene Graphs: Generation and Application
- Discrete and continuous representations and processing in deep learning: Looking forward
- iPerceive: Applying Common-Sense Reasoning to Multi-Modal Dense Video Captioning and Video Question Answering
- Tackling the Unannotated: Scene Graph Generation with Bias-Reduced Models
- Tackling the Challenges in Scene Graph Generation with Local-to-Global Interactions
- Visual Relationship Detection using Scene Graphs: A Survey
- More Grounded Image Captioning by Distilling Image-Text Matching Model
- Towards Causality-Aware Inferring: A Sequential Discriminative Approach for Medical Diagnosis
- Dual ResGCN for Balanced Scene GraphGeneration
- Counterfactual Variable Control for Robust and Interpretable Question Answering
- iReason: Multimodal Commonsense Reasoning using Videos and Natural Language with Interpretability
- Is Object Detection Necessary for Human-Object Interaction Recognition?
- Recovering the Unbiased Scene Graphs from the Biased Ones
- Human-centric Relation Segmentation: Dataset and Solution