1 paper · 1 filter
Shaoyang Cui, Lingbei Meng, Yaodi Luo +1
Video-grounded numerical reasoning requires Vision-Language Models (VLMs) to identify, track, and combine quantitative evidence across frames, actions, and scene changes. Existing…