collaborators

5 papers

cs.CV2026

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation

Massimo Rizzoli, Simone Alghisi, Seyed Mahed Mousavi +1

Performance gains of Vision Language Models (VLMs) obtained by fine-tuning are generally based on ad hoc data collection and annotation of real-world scenes. Despite the improvemen…

cs.CV2026

Getting to the Point: Pointing Improves LVLMs at Counting

Simone Alghisi, Massimo Rizzoli, Seyed Mahed Mousavi +1

Pointing-based methods decompose complex tasks as sequential grounding and reasoning steps. Given a query, the model first grounds the relevant objects by generating their coordina…

cs.AI2026

V-DyKnow: A Dynamic Benchmark for Time-Sensitive Knowledge in Vision Language Models

Seyed Mahed Mousavi, Christian Moiola, Massimo Rizzoli +2

Vision-Language Models (VLMs) are trained on data snapshots of documents, including images and texts. Their training data and evaluation benchmarks are typically static, implicitly…

cs.CV2025

[De|Re]constructing VLMs' Reasoning in Counting

Simone Alghisi, Gabriel Roccabruna, Massimo Rizzoli +2

Vision-Language Models (VLMs) have recently gained attention due to their competitive performance on multiple downstream tasks, achieved by following user-input instructions. Howev…

cs.CV2025

CIVET: Systematic Evaluation of Understanding in VLMs

Massimo Rizzoli, Simone Alghisi, Olha Khomyn +3

While Vision-Language Models (VLMs) have achieved competitive performance in various tasks, their comprehension of the underlying structure and semantics of a scene remains underst…