most citedIntelligent Learning Rate Distribution to reduce Catastrophic Forgetting in Transformers

3 citations · 7 across the 7 of their papers we have counts for

collaborators

7 papers

cs.LG2024

No learning rates needed: Introducing SALSA -- Stable Armijo Line Search Adaptation

Philip Kenneweg, Tristan Kenneweg, Fabian Fumagalli +1

In recent studies, line search methods have been demonstrated to significantly enhance the performance of conventional stochastic gradient descent techniques across various dataset…

cs.CL20243 cited

Intelligent Learning Rate Distribution to reduce Catastrophic Forgetting in Transformers

Philip Kenneweg, Alexander Schulz, Sarah Schröder +1

Pretraining language models on large text corpora is a common practice in natural language processing. Fine-tuning of these models is then performed to achieve the best results on…

cs.CL2024

Debiasing Sentence Embedders through Contrastive Word Pairs

Philip Kenneweg, Sarah Schröder, Alexander Schulz +1

Over the last years, various sentence embedders have been an integral part in the success of current machine learning approaches to Natural Language Processing (NLP). Unfortunately…

cs.AI2024

Neural Architecture Search for Sentence Classification with BERT

Philip Kenneweg, Sarah Schröder, Barbara Hammer

Pre training of language models on large text corpora is common practice in Natural Language Processing. Following, fine tuning of these models is performed to achieve the best res…

cs.LG20241 cited

Improving Line Search Methods for Large Scale Neural Network Training

Philip Kenneweg, Tristan Kenneweg, Barbara Hammer

In recent studies, line search methods have shown significant improvements in the performance of traditional stochastic gradient descent techniques, eliminating the need for a spec…

cs.LG20242 cited

Faster Convergence for Transformer Fine-tuning with Line Search Methods

Philip Kenneweg, Leonardo Galli, Tristan Kenneweg +1

Recent works have shown that line search methods greatly increase performance of traditional stochastic gradient descent methods on a variety of datasets and architectures [1], [2]…