paper

QIRL: Optimized Question-Image Relation Learning for Bias-Robust Visual Question Answering

arXiv:2504.03337

Abstract

Existing bias mitigation methods for Visual Question Answering (VQA), a typical Artificial intelligence application, endure two main limitations. First, they fail to capture the optimal relation between images and texts, as prevailing learning frameworks lack the capacity to extract deep correlations from highly contrasting samples. Second, they overlook assessing Question-Image (QI) relevance during inference, since prior work has not examined the degree of input relevance in debiasing studies. To address these issues, we propose a novel neural network framework termed Optimized Question-Image Relation Learning (QIRL), which provides a reliable implementation of artificial intelligence for VQA tasks and improves the robustness of conventional VQA models through a generation-driven self-supervised learning strategy. Specifically, two modules are introduced. The Negative Image Generation (NIG) module automatically produces highly irrelevant QI pairs during training to strengthen relational learning. In contrast, the Irrelevant Sample Identification (ISI) module enhances model robustness by detecting and filtering out irrelevant inputs, thereby reducing prediction errors. Moreover, to verify the effectiveness of filtering out unrelated QI pairs in mitigating output errors, we propose a specialized metric to evaluate the ISI module's performance. Notably, our approach is model-agnostic and can be seamlessly integrated with various VQA architectures. Extensive experiments on VQA-CPv2 and VQA-v2 datasets demonstrate the effectiveness and generalization ability of our method. Our approach achieves state-of-the-art performance.

QIRL: Optimized Question-Image Relation Learning for Bias-Robust Visual Question Answering · wovepaper