2 citations · 2 across the 2 of their papers we have counts for
3 papers · 1 filter
With a Little Help from your own Past: Prototypical Memory Networks for Image Captioning
Manuele Barraco, Sara Sarto, Marcella Cornia +2
Image captioning, like many tasks involving vision and language, currently relies on Transformer-based architectures for extracting the semantics in an image and translating it int…
Positive-Augmented Contrastive Learning for Image and Video Captioning Evaluation
Sara Sarto, Manuele Barraco, Marcella Cornia +2
The CLIP model has been recently proven to be very effective for a variety of cross-modal tasks, including the evaluation of captions generated from vision-and-language architectur…
CaMEL: Mean Teacher Learning for Image Captioning
Manuele Barraco, Matteo Stefanini, Marcella Cornia +3
Describing images in natural language is a fundamental step towards the automatic modeling of connections between the visual and textual modalities. In this paper we present CaMEL,…