3 papers
cs.CL2026
GLUScope: A Tool for Analyzing GLU Neurons in Transformer Language Models
Sebastian Gerstner, Hinrich Schütze
We present GLUScope, an open-source tool for analyzing neurons in Transformer-based language models, intended for interpretability researchers. We focus on more recent models than…
cs.CL2025
BlackboxNLP-2025 MIB Shared Task: Exploring Ensemble Strategies for Circuit Localization Methods
Philipp Mondorf, Mingyang Wang, Sebastian Gerstner +6
The Circuit Localization track of the Mechanistic Interpretability Benchmark (MIB) evaluates methods for localizing circuits within large language models (LLMs), i.e., subnetworks…
cs.LG2025
Understanding Gated Neurons in Transformers from Their Input-Output Functionality
Sebastian Gerstner, Hinrich Schütze
Interpretability researchers have attempted to understand MLP neurons of language models based on both the contexts in which they activate and their output weight vectors. They hav…