1 paper
Shiho Matta, Lis Kanashiro Pereira, Peitao Han +2
Modern vision-language models (VLMs) excel at many multimodal tasks, yet their grasp of temporal information in video remains weak and has not been adequately evaluated. We probe t…