collaborators

6 papers

cs.CV2026

Visual Instruction Tuning Aligns Modalities through Abstraction

Luis Palacios, Lorenzo Basile, Diego Doimo +1

Visual instruction tuning effectively adapts a pre-trained Large Language Model (LLM) to process image information alongside text. Yet, it remains unclear how visual features are e…

cs.CV2026

When Seeing Overrides Knowing: Disentangling Knowledge Conflicts in Vision-Language Models

Francesco Ortu, Zhijing Jin, Diego Doimo +1

Vision-language models (VLMs) increasingly combine visual and textual information to perform complex tasks. However, conflicts between their internal knowledge and external visual…

cs.CV2026

The Narrow Gate: Localized Image-Text Communication in Native Multimodal Models

Alessandro Pietro Serra, Francesco Ortu, Emanuele Panizon +5

Recent advances in multimodal training have significantly improved the integration of image understanding and generation within a unified model. This study investigates how vision-…

cs.CV2026

Head Pursuit: Probing Attention Specialization in Multimodal Transformers

Lorenzo Basile, Valentino Maiorca, Diego Doimo +2

Language and vision-language models have shown impressive performance across a wide range of tasks, but their internal mechanisms remain only partly understood. In this work, we st…

cs.LG2025

An unsupervised tour through the hidden pathways of deep neural networks

Diego Doimo

The goal of this thesis is to improve our understanding of the internal mechanisms by which deep artificial neural networks create meaningful representations and are able to genera…

cs.CL2025

Emergence of a High-Dimensional Abstraction Phase in Language Transformers

Emily Cheng, Diego Doimo, Corentin Kervadec +4

A language model (LM) is a mapping from a linguistic context to an output token. However, much remains to be known about this mapping, including how its geometric properties relate…