VALSE: A Task-Independent Benchmark for Vision and Language Models Centered on Linguistic Phenomena
arXiv:2112.07566 · doi:10.18653/v1/2022.acl-long.567
Abstract
We propose VALSE (Vision And Language Structured Evaluation), a novel benchmark designed for testing general-purpose pretrained vision and language (V&L) models for their visio-linguistic grounding capabilities on specific linguistic phenomena. VALSE offers a suite of six tests covering various linguistic constructs. Solving these requires models to ground linguistic phenomena in the visual modality, allowing more fine-grained evaluations than hitherto possible. We build VALSE using methods that support the construction of valid foils, and report results from evaluating five widely-used V&L models. Our experiments suggest that current models have considerable difficulty addressing most phenomena. Hence, we expect VALSE to serve as an important benchmark to measure future progress of pretrained V&L models from a linguistic perspective, complementing the canonical task-centred V&L evaluations.
Paper accepted for publication at ACL 2022 Main; 28 pages, 4 figures, 11 tables
References in corpus (3)
Cited by in corpus (6)
- MM-SHAP: A Performance-agnostic Metric for Measuring Multimodal Contributions in Vision and Language Models & Tasks
- Interpreting Vision and Language Generative Models with Semantic Visual Priors
- BOK-VQA: Bilingual outside Knowledge-Based Visual Question Answering via Graph Representation Pretraining
- Explicitly Representing Syntax Improves Sentence-to-layout Prediction of Unexpected Situations
- Natural Language Processing RELIES on Linguistics
- HNC: Leveraging Hard Negative Captions towards Models with Fine-Grained Visual-Linguistic Comprehension Capabilities