most citedIntelligent Learning Rate Distribution to reduce Catastrophic Forgetting in Transformers

3 citations · 3 across the 3 of their papers we have counts for

collaborators

5 papers

cs.CL20243 cited

Intelligent Learning Rate Distribution to reduce Catastrophic Forgetting in Transformers

Philip Kenneweg, Alexander Schulz, Sarah Schröder +1

Pretraining language models on large text corpora is a common practice in natural language processing. Fine-tuning of these models is then performed to achieve the best results on…

cs.CL2024

Debiasing Sentence Embedders through Contrastive Word Pairs

Philip Kenneweg, Sarah Schröder, Alexander Schulz +1

Over the last years, various sentence embedders have been an integral part in the success of current machine learning approaches to Natural Language Processing (NLP). Unfortunately…

cs.AI2024

Neural Architecture Search for Sentence Classification with BERT

Philip Kenneweg, Sarah Schröder, Barbara Hammer

Pre training of language models on large text corpora is common practice in Natural Language Processing. Following, fine tuning of these models is performed to achieve the best res…

cs.LG2024

Targeted Visualization of the Backbone of Encoder LLMs

Isaac Roberts, Alexander Schulz, Luca Hermes +1

Attention based Large Language Models (LLMs) are the state-of-the-art in natural language processing (NLP). The two most common architectures are encoders such as BERT, and decoder…

cs.CL2024

Semantic Properties of cosine based bias scores for word embeddings

Sarah Schröder, Alexander Schulz, Fabian Hinder +1

Plenty of works have brought social biases in language models to attention and proposed methods to detect such biases. As a result, the literature contains a great deal of differen…