4 papers
The GRADIEND Python Package: An End-to-End System for Gradient-Based Feature Learning
Jonathan Drechsel, Steffen Herbold
We present gradiend, an open-source Python package that operationalizes the GRADIEND method for learning feature directions from factual-counterfactual MLM and CLM gradients in lan…
Understanding or Memorizing? A Case Study of German Definite Articles in Language Models
Jonathan Drechsel, Erisa Bytyqi, Steffen Herbold
Language models perform well on grammatical agreement, but it is unclear whether this reflects rule-based generalization or memorization. We study this question for German definite…
GRADIEND: Feature Learning within Neural Networks Exemplified through Biases
Jonathan Drechsel, Steffen Herbold
AI systems frequently exhibit and amplify social biases, leading to harmful consequences in critical areas. This study introduces a novel encoder-decoder approach that leverages mo…
MAMUT: A Novel Framework for Modifying Mathematical Formulas for the Generation of Specialized Datasets for Language Model Training
Jonathan Drechsel, Anja Reusch, Steffen Herbold
Mathematical formulas are a fundamental and widely used component in various scientific fields, serving as a universal language for expressing complex concepts and relationships. W…