1 paper · 1 filter
Suyash Mishra, Qiang Li, Srikanth Patil +2
Vision Language Models (VLMs) have shown strong performance on multimodal reasoning tasks, yet most evaluations focus on short videos and assume unconstrained computational resourc…