Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
PushupBench: Your VLM is not good at counting pushups
Shengzhi Li, Jiarun Chen, Karun Sharma +2
Large vision-language models (VLMs) can recognize \textit{what} happens in video but fail to count \textit{how many} times. We introduce \textbf{PushupBench}, 446 long-form clips (…
cs.CV2024
TARN-VIST: Topic Aware Reinforcement Network for Visual Storytelling
Weiran Chen, Xin Li, Jiaqi Su +4
As a cross-modal task, visual storytelling aims to generate a story for an ordered image sequence automatically. Different from the image captioning task, visual storytelling requi…