1 paper
Shengzhi Li, Jiarun Chen, Karun Sharma +2
Large vision-language models (VLMs) can recognize \textit{what} happens in video but fail to count \textit{how many} times. We introduce \textbf{PushupBench}, 446 long-form clips (…