1 paper
Ivana BeÅová, Michal Gregor, Albert Gatt
How do vision-language (VL) transformer models ground verb phrases and do they integrate contextual and world knowledge in this process? We introduce the CV-Probes dataset, contain…