1 paper
Manuele Barraco, Sara Sarto, Marcella Cornia +2
Image captioning, like many tasks involving vision and language, currently relies on Transformer-based architectures for extracting the semantics in an image and translating it int…