55 citations · 81 across the 8 of their papers we have counts for
3 papers · 1 filter
VALSE: A Task-Independent Benchmark for Vision and Language Models Centered on Linguistic Phenomena
Letitia Parcalabescu, Michele Cafagna, Lilitta Muradjan +3
We propose VALSE (Vision And Language Structured Evaluation), a novel benchmark designed for testing general-purpose pretrained vision and language (V&L) models for their visio-lin…
MAGMA -- Multimodal Augmentation of Generative Models through Adapter-based Finetuning
Constantin Eichenberg, Sidney Black, Samuel Weinbach +2
Large-scale pretraining is fast becoming the norm in Vision-Language (VL) modeling. However, prevailing VL approaches are limited by the requirement for labeled data and the use of…
What is Multimodality?
Letitia Parcalabescu, Nils Trost, Anette Frank
The last years have shown rapid developments in the field of multimodal machine learning, combining e.g., vision, text or speech. In this position paper we explain how the field us…