1 paper · 1 filter
Minchan Kwon, Hyounguk Shon, Junmo Kim
Large multimodal models (LMMs) have recently demonstrated remarkable performance in video question answering (VideoQA), yet reasoning over video remains challenging due to high inf…