1 paper
Sahil Shah, S P Sharan, Harsh Goel +4
Video understanding benchmarks have long centered on single-camera settings, where modern multi-modal language models achieve strong performance across image and video tasks. Yet,…