1 paper
Jack Merullo, Srihita Vatsavaya, Lucius Bushnaq +1
We characterize how memorization is represented in transformer models and show that it can be disentangled in the weights of both language models (LMs) and vision transformers (ViT…