23 citations · 49 across the 6 of their papers we have counts for
3 papers · 1 filter
RUBi: Reducing Unimodal Biases in Visual Question Answering
Remi Cadene, Corentin Dancette, Hedi Ben-younes +2
Visual Question Answering (VQA) is the task of answering questions about an image. Some VQA models often exploit unimodal biases to provide the correct answer without using the ima…
MUREL: Multimodal Relational Reasoning for Visual Question Answering
Remi Cadene, Hedi Ben-younes, Matthieu Cord +1
Multimodal attentional networks are currently state-of-the-art models for Visual Question Answering (VQA) tasks involving real images. Although attention allows to focus on the vis…
BLOCK: Bilinear Superdiagonal Fusion for Visual Question Answering and Visual Relationship Detection
Hedi Ben-younes, Rémi Cadene, Nicolas Thome +1
Multimodal representation learning is gaining more and more interest within the deep learning community. While bilinear models provide an interesting framework to find subtle combi…