1 paper · 1 filter
Marek Hradil, Danae Sánchez Villegas
Vision-language models (VLMs) achieve strong performance on video and image-sequence benchmarks, yet it remains unclear whether they capture temporal structure. To study this quest…