2 citations · 2 across the 11 of their papers we have counts for
4 papers · 1 filter
Recursive Scaling in Masked Diffusion Models
Alba Carballo-Castro, Julianna Piskorz, Paulius Rauba +2
Masked diffusion models (MDMs) have recently emerged as a promising paradigm for sequence generation. Scaling MDMs is conventionally achieved by increasing the parameter count or t…
Task Addition and Weight Disentanglement in Closed-Vocabulary Models
Adam Hazimeh, Alessandro Favero, Pascal Frossard
Task arithmetic has recently emerged as a promising method for editing pre-trained \textit{open-vocabulary} models, offering a cost-effective alternative to standard multi-task fin…
Backdoor Unlearning by Linear Task Decomposition
Amel Abdelraheem, Alessandro Favero, Gerome Bovet +1
Foundation models have revolutionized computer vision by enabling broad generalization across diverse tasks. Yet, they remain highly susceptible to adversarial perturbations and ta…
LiNeS: Post-training Layer Scaling Prevents Forgetting and Enhances Model Merging
Ke Wang, Nikolaos Dimitriadis, Alessandro Favero +3
Fine-tuning pre-trained models has become the standard approach to endow them with specialized knowledge, but it poses fundamental challenges. In particular, \textit{(i)} fine-tuni…