10 citations · 11 across the 4 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2023
Interpretable Visual Question Answering via Reasoning Supervision
Maria Parelli, Dimitrios Mallis, Markos Diomataris +1
Transformer-based architectures have recently demonstrated remarkable performance in the Visual Question Answering (VQA) task. However, such models are likely to disregard crucial…
cs.CV2019
Deeply Supervised Multimodal Attentional Translation Embeddings for Visual Relationship Detection
Nikolaos Gkanatsios, Vassilis Pitsikalis, Petros Koutras +2
Detecting visual relationships, i.e. <Subject, Predicate, Object> triplets, is a challenging Scene Understanding task approached in the past via linguistic priors or spatial inform…