4 papers · 1 filter
Visual Instruction Tuning Aligns Modalities through Abstraction
Luis Palacios, Lorenzo Basile, Diego Doimo +1
Visual instruction tuning effectively adapts a pre-trained Large Language Model (LLM) to process image information alongside text. Yet, it remains unclear how visual features are e…
Head Pursuit: Probing Attention Specialization in Multimodal Transformers
Lorenzo Basile, Valentino Maiorca, Diego Doimo +2
Language and vision-language models have shown impressive performance across a wide range of tasks, but their internal mechanisms remain only partly understood. In this work, we st…
The Narrow Gate: Localized Image-Text Communication in Native Multimodal Models
Alessandro Pietro Serra, Francesco Ortu, Emanuele Panizon +5
Recent advances in multimodal training have significantly improved the integration of image understanding and generation within a unified model. This study investigates how vision-…
ResiDual Transformer Alignment with Spectral Decomposition
Lorenzo Basile, Valentino Maiorca, Luca Bortolussi +2
When examined through the lens of their residual streams, a puzzling property emerges in transformer networks: residual contributions (e.g., attention heads) sometimes specialize i…