1 paper
Noriyuki Kugo, Xiang Li, Zixin Li +9
Video Question Answering (VQA) inherently relies on multimodal reasoning, integrating visual, temporal, and linguistic cues to achieve a deeper understanding of video content. Howe…