activity
20242026
collaborators

10 papers

cs.CV2026

Visual Instruction Tuning Aligns Modalities through Abstraction

Luis Palacios, Lorenzo Basile, Diego Doimo +1

Visual instruction tuning effectively adapts a pre-trained Large Language Model (LLM) to process image information alongside text. Yet, it remains unclear how visual features are e…

cs.CV2026

When Seeing Overrides Knowing: Disentangling Knowledge Conflicts in Vision-Language Models

Francesco Ortu, Zhijing Jin, Diego Doimo +1

Vision-language models (VLMs) increasingly combine visual and textual information to perform complex tasks. However, conflicts between their internal knowledge and external visual…

cs.CY2026

Preserving Historical Truth: Detecting Historical Revisionism in Large Language Models

Francesco Ortu, Joeun Yook, Punya Syon Pandey +5

Large language models (LLMs) are increasingly used as sources of historical information, motivating the need for scalable audits on contested events and politically charged narrati…

cs.CV2026

The Narrow Gate: Localized Image-Text Communication in Native Multimodal Models

Alessandro Pietro Serra, Francesco Ortu, Emanuele Panizon +5

Recent advances in multimodal training have significantly improved the integration of image understanding and generation within a unified model. This study investigates how vision-…

cs.CV2026

Head Pursuit: Probing Attention Specialization in Multimodal Transformers

Lorenzo Basile, Valentino Maiorca, Diego Doimo +2

Language and vision-language models have shown impressive performance across a wide range of tasks, but their internal mechanisms remain only partly understood. In this work, we st…

cs.CL2025

Are LLMs Good Safety Agents or a Propaganda Engine?

Neemesh Yadav, Francesco Ortu, Jiarui Liu +5

Large Language Models (LLMs) are trained to refuse to respond to harmful content. However, systematic analyses of whether this behavior is truly a reflection of its safety policies…