71 citations · 90 across the 5 of their papers we have counts for
1 paper · 1 filter
Yonatan Bitton, Nitzan Bitton Guetta, Ron Yosef +4
While vision-and-language models perform well on tasks such as visual question answering, they struggle when it comes to basic human commonsense reasoning skills. In this work, we…