activity
20162025
most citedRetrieval-Augmented Transformer for Image Captioning

9 citations · 10 across the 7 of their papers we have counts for

collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV20229 cited

Retrieval-Augmented Transformer for Image Captioning

Sara Sarto, Marcella Cornia, Lorenzo Baraldi +1

Image captioning models aim at connecting Vision and Language by providing natural language descriptions of input images. In the past few years, the task has been tackled by learni…

cs.CV2022

Boosting Modern and Historical Handwritten Text Recognition with Deformable Convolutions

Silvia Cascianelli, Marcella Cornia, Lorenzo Baraldi +1

Handwritten Text Recognition (HTR) in free-layout pages is a challenging image understanding task that can provide a relevant boost to the digitization of handwritten documents and…

cs.CV2022

The LAM Dataset: A Novel Benchmark for Line-Level Handwritten Text Recognition

Silvia Cascianelli, Vittorio Pippi, Martin Maarand +4

Handwritten Text Recognition (HTR) is an open problem at the intersection of Computer Vision and Natural Language Processing. The main challenges, when dealing with historical manu…

cs.CV2022

ALADIN: Distilling Fine-grained Alignment Scores for Efficient Image-Text Matching and Retrieval

Nicola Messina, Matteo Stefanini, Marcella Cornia +4

Image-text matching is gaining a leading role among tasks involving the joint understanding of vision and language. In literature, this task is often used as a pre-training objecti…

cs.CV20161 cited

Recognizing and Presenting the Storytelling Video Structure with Deep Multimodal Networks

Lorenzo Baraldi, Costantino Grana, Rita Cucchiara

This paper presents a novel approach for temporal and semantic segmentation of edited videos into meaningful segments, from the point of view of the storytelling structure. The obj…