10 citations · 22 across the 6 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2019
What is needed for simple spatial language capabilities in VQA?
Alexander Kuhnle, Ann Copestake
Visual question answering (VQA) comprises a variety of language capabilities. The diagnostic benchmark dataset CLEVR has fueled progress by helping to better assess and distinguish…
cs.CV2018
The meaning of "most" for visual question answering models
Alexander Kuhnle, Ann Copestake
The correct interpretation of quantifier statements in the context of a visual scene requires non-trivial inference mechanisms. For the example of "most", we discuss two strategies…