1 paper · 1 filter
Yu Han, Wenhao Li, Yichao Cao +4
Existing VLMs have achieved strong performance in video understanding, yet they struggle with long-video spatiotemporal reasoning when target objects become invisible, often mistak…