2 papers
cs.LG2025
MINERVA: Evaluating Complex Video Reasoning
Arsha Nagrani, Sachit Menon, Ahmet Iscen +9
Multimodal LLMs are turning their focus to video benchmarks, however most video benchmarks only provide outcome supervision, with no intermediate or interpretable reasoning steps.…
cs.LG2025
Neptune: The Long Orbit to Benchmarking Long Video Understanding
Arsha Nagrani, Mingda Zhang, Ramin Mehran +10
We introduce Neptune, a benchmark for long video understanding that requires reasoning over long time horizons and across different modalities. Many existing video datasets and mod…