8 citations · 10 across the 4 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2022
Understanding Cross-modal Interactions in V&L Models that Generate Scene Descriptions
Michele Cafagna, Kees van Deemter, Albert Gatt
Image captioning models tend to describe images in an object-centric way, emphasising visible objects. But image descriptions can also abstract away from objects and describe the t…
cs.CL2021★ 8 cited
What Vision-Language Models `See' when they See Scenes
Michele Cafagna, Kees van Deemter, Albert Gatt
Images can be described in terms of the objects they contain, or in terms of the types of scene or place that they instantiate. In this paper we address to what extent pretrained V…