1 paper
Man Luo, Shailaja Keyur Sampat, Riley Tallman +4
GQA~\citep{hudson2019gqa} is a dataset for real-world visual reasoning and compositional question answering. We found that many answers predicted by the best vision-language models…