1 paper
Kun Li, Michael Ying Yang, Sami Sebastian Brandt
Audio--Visual Question Answering (AVQA) is a challenging multimodal task that requires jointly reasoning over audio, visual, and textual information in a given video to answer natu…