1 paper · 1 filter
Haodi Ma, Vyom Pathak, Daisy Zhe Wang
Video Question Answering (VQA) requires models to reason over spatial, temporal, and causal cues in videos. Recent vision language models (VLMs) achieve strong results but often re…