5 papers
Judge Circuits
Nils Feldhus, Tanja Baeumel, Elena Golimblevskaia +10
LLM-as-a-judge has become the dominant paradigm for grading model outputs at scale, yet the same model assigns systematically different scores when its output format changes (e.g.,…
Multilingual Steering by Design: Multilingual Sparse Autoencoders and Principled Layer Selection
Yusser Al Ghussin, Daniil Gurgurov, Tanja Baeumel +3
Sparse autoencoders (SAEs) enable feature-level mechanistic interpretability and activation steering in large language models (LLMs), but SAE-based language control remains unrelia…
Disentangling Mathematical Reasoning in LLMs: A Methodological Investigation of Internal Mechanisms
Tanja Baeumel, Josef van Genabith, Simon Ostermann
Large language models (LLMs) have demonstrated impressive capabilities, yet their internal mechanisms for handling reasoning-intensive tasks remain underexplored. To advance the un…
From Weights to Activations: Is Steering the Next Frontier of Adaptation?
Simon Ostermann, Daniil Gurgurov, Tanja Baeumel +4
Post-training adaptation of language models is commonly achieved through parameter updates or input-based methods such as fine-tuning, parameter-efficient adaptation, and prompting…
CLaS-Bench: A Cross-Lingual Alignment and Steering Benchmark
Daniil Gurgurov, Yusser Al Ghussin, Tanja Baeumel +5
Understanding and controlling the behavior of large language models (LLMs) is an increasingly important topic in multilingual NLP. Beyond prompting or fine-tuning, , i.e.,~manipula…