1 citations · 2 across the 5 of their papers we have counts for
1 paper · 2 filters
Hongyeob Kim, Inyoung Jung, Dayoon Suh +3
Audio-Visual Question Answering (AVQA) requires not only question-based multimodal reasoning but also precise temporal grounding to capture subtle dynamics for accurate prediction.…