1 paper · 1 filter
Zihang Fu, Haonan Wang, Jian Kang +2
Multimodal adaptation can erode temporal reasoning (TR) in video-language models (VLMs), leaving models able to perceive salient events yet unable to infer their temporal and causa…