From the 1 of 6 linked papers with an AI index.
6 papers
Attentive multilayer fusion for vision transformers
Laure Ciernik, Marco Morik, Lukas Thede +4
The paper introduces Attentive Layer Fusion (ALF), a method that dynamically combines representations from all layers of a Vision Transformer to improve linear probing on downstrea…
CapTrack: Multifaceted Evaluation of Forgetting in LLM Post-Training
Lukas Thede, Stefan Winzeck, Zeynep Akata +1
Large language model (LLM) post-training enhances latent skills, unlocks value alignment, improves performance, and enables domain adaptation. Unfortunately, post-training is known…
Structural Pruning of Large Vision Language Models: A Comprehensive Study on Pruning Dynamics, Recovery, and Data Efficiency
Yiran Huang, Lukas Thede, Massimiliano Mancini +2
While Large Vision Language Models (LVLMs) demonstrate impressive capabilities, their substantial computational and memory requirements pose deployment challenges on resource-const…
WikiBigEdit: Understanding the Limits of Lifelong Knowledge Editing in LLMs
Lukas Thede, Karsten Roth, Matthias Bethge +2
Keeping large language models factually up-to-date is crucial for deployment, yet costly retraining remains a challenge. Knowledge editing offers a promising alternative, but metho…
Reflecting on the State of Rehearsal-free Continual Learning with Pretrained Models
Lukas Thede, Karsten Roth, Olivier J. Hénaff +2
With the advent and recent ubiquity of foundation models, continual learning (CL) has recently shifted from continual training from scratch to the continual adaptation of pretraine…
Investigating Structural Pruning and Recovery Techniques for Compressing Multimodal Large Language Models: An Empirical Study
Yiran Huang, Lukas Thede, Massimiliano Mancini +2
While Multimodal Large Language Models (MLLMs) demonstrate impressive capabilities, their substantial computational and memory requirements pose significant barriers to practical d…