8 papers
Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation
Massimo Rizzoli, Simone Alghisi, Seyed Mahed Mousavi +1
Performance gains of Vision Language Models (VLMs) obtained by fine-tuning are generally based on ad hoc data collection and annotation of real-world scenes. Despite the improvemen…
Getting to the Point: Pointing Improves LVLMs at Counting
Simone Alghisi, Massimo Rizzoli, Seyed Mahed Mousavi +1
Pointing-based methods decompose complex tasks as sequential grounding and reasoning steps. Given a query, the model first grounds the relevant objects by generating their coordina…
V-DyKnow: A Dynamic Benchmark for Time-Sensitive Knowledge in Vision Language Models
Seyed Mahed Mousavi, Christian Moiola, Massimo Rizzoli +2
Vision-Language Models (VLMs) are trained on data snapshots of documents, including images and texts. Their training data and evaluation benchmarks are typically static, implicitly…
LLMs as Repositories of Factual Knowledge: Limitations and Solutions
Seyed Mahed Mousavi, Simone Alghisi, Giuseppe Riccardi
LLMs' sources of knowledge are data snapshots containing factual information about entities collected at different timestamps and from different media types (e.g. wikis, social med…
What Does Loss Optimization Actually Teach, If Anything? Knowledge Dynamics in Continual Pre-training of LLMs
Seyed Mahed Mousavi, Simone Alghisi, Giuseppe Riccardi
Continual Pre-Training (CPT) is widely used for acquiring and updating factual knowledge in LLMs. This practice treats loss as a proxy for knowledge learning, while offering no gro…
[De|Re]constructing VLMs' Reasoning in Counting
Simone Alghisi, Gabriel Roccabruna, Massimo Rizzoli +2
Vision-Language Models (VLMs) have recently gained attention due to their competitive performance on multiple downstream tasks, achieved by following user-input instructions. Howev…