8 citations · 9 across the 2 of their papers we have counts for
2 papers
cs.CL2021★ 8 cited
What Vision-Language Models `See' when they See Scenes
Michele Cafagna, Kees van Deemter, Albert Gatt
Images can be described in terms of the objects they contain, or in terms of the types of scene or place that they instantiate. In this paper we address to what extent pretrained V…
cs.NE2019★ 1 cited
On Architectures for Including Visual Information in Neural Language Models for Image Description
Marc Tanti, Albert Gatt, Kenneth P. Camilleri
A neural language model can be conditioned into generating descriptions for images by providing visual information apart from the sentence prefix. This visual information can be in…