1 paper
Guangyao Li, Henghui Du, Di Hu
The Audio Visual Question Answering (AVQA) task aims to answer questions related to various visual objects, sounds, and their interactions in videos. Such naturally multimodal vide…