3 papers
cs.CL2026
GKnow: Measuring the Entanglement of Gender Bias and Factual Gender
Leonor Veloso, Hinrich Schütze
Recent works have analyzed the impact of individual components of neural networks on gendered predictions, often with a focus on mitigating gender bias. However, mechanistic interp…
cs.CL2025
BlackboxNLP-2025 MIB Shared Task: Exploring Ensemble Strategies for Circuit Localization Methods
Philipp Mondorf, Mingyang Wang, Sebastian Gerstner +6
The Circuit Localization track of the Mechanistic Interpretability Benchmark (MIB) evaluates methods for localizing circuits within large language models (LLMs), i.e., subnetworks…
cs.CL2025
SLAyiNG: A Diverse and Community-validated Dataset of Queer Slang
Leonor Veloso, Lea Hirlimann, Philipp Wicke +4
Queer vernacular is rarely studied in NLP, despite advancements in resources and evaluation for other sociolects and informal language. Because of this, NLP systems often process q…