11 papers
Learn from your own latents and not from tokens: A sample-complexity theory
Daniel J. Korchinski, Alessandro Favero, Matthieu Wyart
Generative models, from diffusion models to large language models, achieve remarkable performance but at a cost in training data orders of magnitude larger than what biological lea…
MEMOIR: Lifelong Model Editing with Minimal Overwrite and Informed Retention for LLMs
Ke Wang, Yiming Qin, Nikolaos Dimitriadis +2
Language models deployed in real-world systems often require post-hoc updates to incorporate new or corrected knowledge. However, editing such models efficiently and reliably-witho…
Task Addition and Weight Disentanglement in Closed-Vocabulary Models
Adam Hazimeh, Alessandro Favero, Pascal Frossard
Task arithmetic has recently emerged as a promising method for editing pre-trained \textit{open-vocabulary} models, offering a cost-effective alternative to standard multi-task fin…
Backdoor Unlearning by Linear Task Decomposition
Amel Abdelraheem, Alessandro Favero, Gerome Bovet +1
Foundation models have revolutionized computer vision by enabling broad generalization across diverse tasks. Yet, they remain highly susceptible to adversarial perturbations and ta…
The Physics of Data and Tasks: Theories of Locality and Compositionality in Deep Learning
Alessandro Favero
Deep neural networks have achieved remarkable success, yet our understanding of how they learn remains limited. These models can learn high-dimensional tasks, which is generally st…
Bigger Isn't Always Memorizing: Early Stopping Overparameterized Diffusion Models
Alessandro Favero, Antonio Sclocchi, Matthieu Wyart
Diffusion probabilistic models have become a cornerstone of modern generative AI, yet the mechanisms underlying their generalization remain poorly understood. In fact, if these mod…