6 papers
Visual Instruction Tuning Aligns Modalities through Abstraction
Luis Palacios, Lorenzo Basile, Diego Doimo +1
Visual instruction tuning effectively adapts a pre-trained Large Language Model (LLM) to process image information alongside text. Yet, it remains unclear how visual features are e…
When Seeing Overrides Knowing: Disentangling Knowledge Conflicts in Vision-Language Models
Francesco Ortu, Zhijing Jin, Diego Doimo +1
Vision-language models (VLMs) increasingly combine visual and textual information to perform complex tasks. However, conflicts between their internal knowledge and external visual…
The Narrow Gate: Localized Image-Text Communication in Native Multimodal Models
Alessandro Pietro Serra, Francesco Ortu, Emanuele Panizon +5
Recent advances in multimodal training have significantly improved the integration of image understanding and generation within a unified model. This study investigates how vision-…
Head Pursuit: Probing Attention Specialization in Multimodal Transformers
Lorenzo Basile, Valentino Maiorca, Diego Doimo +2
Language and vision-language models have shown impressive performance across a wide range of tasks, but their internal mechanisms remain only partly understood. In this work, we st…
An unsupervised tour through the hidden pathways of deep neural networks
Diego Doimo
The goal of this thesis is to improve our understanding of the internal mechanisms by which deep artificial neural networks create meaningful representations and are able to genera…
Emergence of a High-Dimensional Abstraction Phase in Language Transformers
Emily Cheng, Diego Doimo, Corentin Kervadec +4
A language model (LM) is a mapping from a linguistic context to an output token. However, much remains to be known about this mapping, including how its geometric properties relate…