1 paper
Shaden Shaar, Bradon Thymes, Sirawut Chaixanien +2
Understanding real-world videos such as movies requires integrating visual and dialogue cues. Yet existing VideoQA benchmarks struggle to capture this multimodal reasoning and, giv…