57 citations · 70 across the 6 of their papers we have counts for
1 paper · 1 filter
Arsha Nagrani, Sachit Menon, Ahmet Iscen +9
Multimodal LLMs are turning their focus to video benchmarks, however most video benchmarks only provide outcome supervision, with no intermediate or interpretable reasoning steps.…