1 paper
Suryanarayana Reddy Yarrabothula, Manisha Chawla, Kunal Sinha +5
Existing video benchmarks evaluate action recognition on consumer videos, egocentric recordings, or simulated industrial environments. They do not test vision-language models under…