activity
20172022
most citedTransfer learning from language models to image caption generators: Better models may not transfer better

3 citations · 5 across the 3 of their papers we have counts for

collaborators

9 papers

cs.CV20221 cited

Face2Text revisited: Improved data set and baseline results

Marc Tanti, Shaun Abdilla, Adrian Muscat +3

Current image description generation models do not transfer well to the task of describing human faces. To encourage the development of more human-focused descriptions, we develope…

cs.NE20191 cited

On Architectures for Including Visual Information in Neural Language Models for Image Description

Marc Tanti, Albert Gatt, Kenneth P. Camilleri

A neural language model can be conditioned into generating descriptions for images by providing visual information apart from the sentence prefix. This visual information can be in…

cs.CL2019

Visuallly Grounded Generation of Entailments from Premises

Somaye Jafaritazehjani, Albert Gatt, Marc Tanti

Natural Language Inference (NLI) is the task of determining the semantic relationship between a premise and a hypothesis. In this paper, we focus on the {\em generation} of hypothe…

cs.CL20193 cited

Transfer learning from language models to image caption generators: Better models may not transfer better

Marc Tanti, Albert Gatt, Kenneth P. Camilleri

When designing a neural caption generator, a convolutional neural network can be used to extract image features. Is it possible to also use a neural language model to extract sente…

cs.NE2018

Quantifying the amount of visual information used by neural caption generators

Marc Tanti, Albert Gatt, Kenneth P. Camilleri

This paper addresses the sensitivity of neural image caption generators to their visual input. A sensitivity analysis and omission analysis based on image foils is reported, showin…

cs.NE2018

Pre-gen metrics: Predicting caption quality metrics without generating captions

Marc Tanti, Albert Gatt, Adrian Muscat

Image caption generation systems are typically evaluated against reference outputs. We show that it is possible to predict output quality without generating the captions, based on…