1 citations · 1 across the 1 of their papers we have counts for
1 paper · 1 filter
Arsha Nagrani, Sachit Menon, Ahmet Iscen +9
Multimodal LLMs are turning their focus to video benchmarks, however most video benchmarks only provide outcome supervision, with no intermediate or interpretable reasoning steps.…