2 papers
cs.CV2023
Self-Supervised Learning for Visual Relationship Detection through Masked Bounding Box Reconstruction
Zacharias Anastasakis, Dimitrios Mallis, Markos Diomataris +3
We present a novel self-supervised approach for representation learning, particularly for the task of Visual Relationship Detection (VRD). Motivated by the effectiveness of Masked…
cs.CV2023
Interpretable Visual Question Answering via Reasoning Supervision
Maria Parelli, Dimitrios Mallis, Markos Diomataris +1
Transformer-based architectures have recently demonstrated remarkable performance in the Visual Question Answering (VQA) task. However, such models are likely to disregard crucial…