1 paper
Siwen Luo, Soyeon Caren Han, Kaiyuan Sun +1
Visual question answering (VQA) is a challenging multi-modal task that requires not only the semantic understanding of both images and questions, but also the sound perception of a…