7 papers
The Narrow Gate: Localized Image-Text Communication in Native Multimodal Models
Alessandro Pietro Serra, Francesco Ortu, Emanuele Panizon +5
Recent advances in multimodal training have significantly improved the integration of image understanding and generation within a unified model. This study investigates how vision-…
Density-Informed VAE (DiVAE): Reliable Log-Prior Probability via Density Alignment Regularization
Michele Alessi, Alessio Ansuini, Alex Rodriguez
We introduce Density-Informed VAE (DiVAE), a lightweight, data-driven regularizer that aligns the VAE log-prior probability with a log-density estimated from data. St…
Persistent Topological Features in Large Language Models
Yuri Gardinazzi, Karthik Viswanathan, Giada Panerai +3
Understanding the decision-making processes of large language models is critical given their widespread applications. To achieve this, we aim to connect a formal mathematical frame…
Emergent representations in networks trained with the Forward-Forward algorithm
Niccolò Tosato, Lorenzo Basile, Emanuele Ballarin +3
The Backpropagation algorithm has often been criticised for its lack of biological realism. In an attempt to find a more biologically plausible alternative, the recently introduced…
Interpreting and Steering Protein Language Models through Sparse Autoencoders
Edith Natalia Villegas Garcia, Alessio Ansuini
The rapid advancements in transformer-based language models have revolutionized natural language processing, yet understanding the internal mechanisms of these models remains a sig…
The representation landscape of few-shot learning and fine-tuning in large language models
Diego Doimo, Alessandro Serra, Alessio Ansuini +1
In-context learning (ICL) and supervised fine-tuning (SFT) are two common strategies for improving the performance of modern large language models (LLMs) on specific tasks. Despite…