11 papers
Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation
Massimo Rizzoli, Simone Alghisi, Seyed Mahed Mousavi +1
Performance gains of Vision Language Models (VLMs) obtained by fine-tuning are generally based on ad hoc data collection and annotation of real-world scenes. Despite the improvemen…
Getting to the Point: Pointing Improves LVLMs at Counting
Simone Alghisi, Massimo Rizzoli, Seyed Mahed Mousavi +1
Pointing-based methods decompose complex tasks as sequential grounding and reasoning steps. Given a query, the model first grounds the relevant objects by generating their coordina…
V-DyKnow: A Dynamic Benchmark for Time-Sensitive Knowledge in Vision Language Models
Seyed Mahed Mousavi, Christian Moiola, Massimo Rizzoli +2
Vision-Language Models (VLMs) are trained on data snapshots of documents, including images and texts. Their training data and evaluation benchmarks are typically static, implicitly…
LLMs as Repositories of Factual Knowledge: Limitations and Solutions
Seyed Mahed Mousavi, Simone Alghisi, Giuseppe Riccardi
LLMs' sources of knowledge are data snapshots containing factual information about entities collected at different timestamps and from different media types (e.g. wikis, social med…
MATEO: A Multimodal Benchmark for Temporal Reasoning and Planning in LVLMs
Gabriel Roccabruna, Olha Khomyn, Giuseppe Riccardi
AI agents need to plan to achieve complex goals that involve orchestrating perception, sub-goal decomposition, and execution. These plans consist of ordered steps structured accord…
What Does Loss Optimization Actually Teach, If Anything? Knowledge Dynamics in Continual Pre-training of LLMs
Seyed Mahed Mousavi, Simone Alghisi, Giuseppe Riccardi
Continual Pre-Training (CPT) is widely used for acquiring and updating factual knowledge in LLMs. This practice treats loss as a proxy for knowledge learning, while offering no gro…