2 citations · 4 across the 5 of their papers we have counts for
10 papers
Two-Scale Latent Dynamics for Recurrent-Depth Transformers
Francesco Pappone, Donato Crisostomi, Emanuele Rodolà
Recurrent-depth transformers scale test-time compute by iterating latent computations before emitting tokens. We study the geometry of these iterates and argue for a simple, two-sc…
On Task Vectors and Gradients
Luca Zhou, Daniele Solombrino, Donato Crisostomi +4
Task arithmetic has emerged as a simple yet powerful technique for model merging, enabling the combination of multiple finetuned models into one. Despite its empirical success, a c…
Implicit Inversion turns CLIP into a Decoder
Antonio D'Orazio, Maria Rosaria Briglia, Donato Crisostomi +3
CLIP is a discriminative model trained to align images and text in a shared embedding space. Due to its multimodal structure, it serves as the backbone of many generative pipelines…
Update Your Transformer to the Latest Release: Re-Basin of Task Vectors
Filippo Rinaldi, Giacomo Capitani, Lorenzo Bonicelli +6
Foundation models serve as the backbone for numerous specialized models developed through fine-tuning. However, when the underlying pretrained model is updated or retrained (e.g.,…
Mergenetic: a Simple Evolutionary Model Merging Library
Adrian Robert Minut, Tommaso Mencattini, Andrea Santilli +2
Model merging allows combining the capabilities of existing models into a new one - post hoc, without additional training. This has made it increasingly popular thanks to its low c…
Activation Patching for Interpretable Steering in Music Generation
Simone Facchiano, Giorgio Strano, Donato Crisostomi +4
Understanding how large audio models represent music, and using that understanding to steer generation, is both challenging and underexplored. Inspired by mechanistic interpretabil…