153 citations · 154 across the 3 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2020
Fine-grained Visual Textual Alignment for Cross-Modal Retrieval using Transformer Encoders
Nicola Messina, Giuseppe Amato, Andrea Esuli +3
Despite the evolution of deep-learning-based visual-textual processing systems, precise multi-modal matching remains a challenging task. In this work, we tackle the task of cross-m…
cs.CV2020
Transformer Reasoning Network for Image-Text Matching and Retrieval
Nicola Messina, Fabrizio Falchi, Andrea Esuli +1
Image-text matching is an interesting and fascinating task in modern AI research. Despite the evolution of deep-learning-based image and text processing systems, multi-modal matchi…