23 citations · 26 across the 2 of their papers we have counts for
8 papers
Beyond Question-Based Biases: Assessing Multimodal Shortcut Learning in Visual Question Answering
Corentin Dancette, Remi Cadene, Damien Teney +1
We introduce an evaluation methodology for visual question answering (VQA) to better diagnose cases of shortcut learning. These cases happen when a model exploits spurious statisti…
Overcoming Statistical Shortcuts for Open-ended Visual Counting
Corentin Dancette, Remi Cadene, Xinlei Chen +1
Machine learning models tend to over-rely on statistical shortcuts. These spurious correlations between parts of the input and the output labels does not hold in real-world setting…
RUBi: Reducing Unimodal Biases in Visual Question Answering
Remi Cadene, Corentin Dancette, Hedi Ben-younes +2
Visual Question Answering (VQA) is the task of answering questions about an image. Some VQA models often exploit unimodal biases to provide the correct answer without using the ima…
MUREL: Multimodal Relational Reasoning for Visual Question Answering
Remi Cadene, Hedi Ben-younes, Matthieu Cord +1
Multimodal attentional networks are currently state-of-the-art models for Visual Question Answering (VQA) tasks involving real images. Although attention allows to focus on the vis…
BLOCK: Bilinear Superdiagonal Fusion for Visual Question Answering and Visual Relationship Detection
Hedi Ben-younes, Rémi Cadene, Nicolas Thome +1
Multimodal representation learning is gaining more and more interest within the deep learning community. While bilinear models provide an interesting framework to find subtle combi…
Images & Recipes: Retrieval in the cooking context
Micael Carvalho, Rémi Cadène, David Picard +2
Recent advances in the machine learning community allowed different use cases to emerge, as its association to domains like cooking which created the computational cuisine. In this…