1 paper
Rahul Jain, Keval Doshi, Burak Uzkent +1
Recent progress in multimodal large language models (MLLMs) has led to a surge of benchmarks for long-video reasoning. However, most existing benchmarks rely on localized cues and…