6 papers
Do LLMs Experience an Internal Polylogue? Investigating Reasoning through the Lens of Personas
Nils A. Herrmann, Leander Girrbach, Kirill Bykov +1
Recent work shows that large language models (LLMs) encode behavioral traits ("personas") as linear directions in activation space, often called "persona vectors". Prior work has u…
Manipulating Feature Visualizations with Gradient Slingshots
Dilyara Bareeva, Marina M. -C. Höhne, Alexander Warnecke +5
Feature Visualization (FV) is a widely used technique for interpreting concepts learned by Deep Neural Networks (DNNs), which synthesizes input patterns that maximally activate a g…
Capturing Polysemanticity with PRISM: A Multi-Concept Feature Description Framework
Laura Kopf, Nils Feldhus, Kirill Bykov +4
Automated interpretability research aims to identify concepts encoded in neural network features to enhance human understanding of model behavior. Within the context of large langu…
Explaining Bayesian Neural Networks
Kirill Bykov, Marina M. -C. Höhne, Adelaida Creosteanu +4
To advance the transparency of learning machines such as Deep Neural Networks (DNNs), the field of Explainable AI (XAI) was established to provide interpretations of DNNs' predicti…
Deep Learning Meets Teleconnections: Improving S2S Predictions for European Winter Weather
Philine L. Bommer, Marlene Kretschmer, Fiona R. Spuler +2
Predictions on subseasonal-to-seasonal (S2S) timescales--ranging from two weeks to two month--are crucial for early warning systems but remain challenging owing to chaos in the cli…
CoSy: Evaluating Textual Explanations of Neurons
Laura Kopf, Philine Lou Bommer, Anna Hedström +3
A crucial aspect of understanding the complex nature of Deep Neural Networks (DNNs) is the ability to explain learned concepts within their latent representations. While methods ex…