90 citations · 182 across the 7 of their papers we have counts for
11 papers · 1 filter
Training Vision-Language Models with Less Bimodal Supervision
Elad Segal, Ben Bogin, Jonathan Berant
Standard practice in pretraining multimodal models, such as vision-language models, is to rely on pairs of aligned inputs from both modalities, for example, aligned image-text pair…
COVR: A test-bed for Visually Grounded Compositional Generalization with real images
Ben Bogin, Shivanshu Gupta, Matt Gardner +1
While interest in models that generalize at test time to new compositions has risen in recent years, benchmarks in the visually-grounded domain have thus far been restricted to syn…
Text-to-SQL in the Wild: A Naturally-Occurring Dataset Based on Stack Exchange Data
Moshe Hazoom, Vibhor Malik, Ben Bogin
Most available semantic parsing datasets, comprising of pairs of natural utterances and logical forms, were collected solely for the purpose of training and evaluation of natural l…
Latent Compositional Representations Improve Systematic Generalization in Grounded Question Answering
Ben Bogin, Sanjay Subramanian, Matt Gardner +1
Answering questions that involve multi-step reasoning requires decomposing them and using the answers of intermediate steps to reach the final answer. However, state-of-the-art mod…
Obtaining Faithful Interpretations from Compositional Neural Networks
Sanjay Subramanian, Ben Bogin, Nitish Gupta +4
Neural module networks (NMNs) are a popular approach for modeling compositionality: they achieve high accuracy when applied to problems in language and vision, while reflecting the…
Evaluating Models' Local Decision Boundaries via Contrast Sets
Matt Gardner, Yoav Artzi, Victoria Basmova +23
Standard test sets for supervised learning evaluate in-distribution generalization. Unfortunately, when a dataset has systematic gaps (e.g., annotation artifacts), these evaluation…