4 citations · 4 across the 8 of their papers we have counts for
Showing 2026Show all
2 papers · 1 filter
cs.AI2026
V-DyKnow: A Dynamic Benchmark for Time-Sensitive Knowledge in Vision Language Models
Seyed Mahed Mousavi, Christian Moiola, Massimo Rizzoli +2
Vision-Language Models (VLMs) are trained on data snapshots of documents, including images and texts. Their training data and evaluation benchmarks are typically static, implicitly…
cs.CV2026
Getting to the Point: Pointing Improves LVLMs at Counting
Simone Alghisi, Massimo Rizzoli, Seyed Mahed Mousavi +1
Pointing-based methods decompose complex tasks as sequential grounding and reasoning steps. Given a query, the model first grounds the relevant objects by generating their coordina…