15 citations · 39 across the 19 of their papers we have counts for
3 papers · 1 filter
Seeing past words: Testing the cross-modal capabilities of pretrained V&L models on counting tasks
Letitia Parcalabescu, Albert Gatt, Anette Frank +1
We investigate the reasoning ability of pretrained vision and language (V&L) models in two tasks that require multimodal integration: (1) discriminating a correct image-sentence pa…
Are scene graphs good enough to improve Image Captioning?
Victor Milewski, Marie-Francine Moens, Iacer Calixto
Many top-performing image captioning models rely solely on object features computed with an object detection model to generate image descriptions. However, recent studies propose t…
ImagiFilter: A resource to enable the semi-automatic mining of images at scale
Houda Alberts, Iacer Calixto
Datasets (semi-)automatically collected from the web can easily scale to millions of entries, but a dataset's usefulness is directly related to how clean and high-quality its examp…