3 papers
cs.LG2026
Grow, Don't Overwrite: Fine-tuning Without Forgetting
Dyah Adila, Hanna Mazzawi, Benoit Dherin +1
Adapting pre-trained models to specialized tasks often leads to catastrophic forgetting, where new knowledge overwrites foundational capabilities. Existing methods either compromis…
cs.LG2024
Majority Kernels: An Approach to Leverage Big Model Dynamics for Efficient Small Model Training
Hanna Mazzawi, Pranjal Awasthi, Xavi Gonzalvo +1
Recent breakthroughs and successful deployment of large language and vision models in a constrained environment predominantly follow a two phase approach. First, large models are t…
cs.LG2024
Deep Fusion: Efficient Network Training via Pre-trained Initializations
Hanna Mazzawi, Xavi Gonzalvo, Michael Wunder +2
In recent years, deep learning has made remarkable progress in a wide range of domains, with a particularly notable impact on natural language processing tasks. One of the challeng…