activity
20242026
collaborators

6 papers

cs.AI2026

Do LLMs Experience an Internal Polylogue? Investigating Reasoning through the Lens of Personas

Nils A. Herrmann, Leander Girrbach, Kirill Bykov +1

Recent work shows that large language models (LLMs) encode behavioral traits ("personas") as linear directions in activation space, often called "persona vectors". Prior work has u…

cs.LG2026

Manipulating Feature Visualizations with Gradient Slingshots

Dilyara Bareeva, Marina M. -C. Höhne, Alexander Warnecke +5

Feature Visualization (FV) is a widely used technique for interpreting concepts learned by Deep Neural Networks (DNNs), which synthesizes input patterns that maximally activate a g…

cs.LG2025

Capturing Polysemanticity with PRISM: A Multi-Concept Feature Description Framework

Laura Kopf, Nils Feldhus, Kirill Bykov +4

Automated interpretability research aims to identify concepts encoded in neural network features to enhance human understanding of model behavior. Within the context of large langu…

cs.LG2025

Explaining Bayesian Neural Networks

Kirill Bykov, Marina M. -C. Höhne, Adelaida Creosteanu +4

To advance the transparency of learning machines such as Deep Neural Networks (DNNs), the field of Explainable AI (XAI) was established to provide interpretations of DNNs' predicti…

cs.LG2025

Deep Learning Meets Teleconnections: Improving S2S Predictions for European Winter Weather

Philine L. Bommer, Marlene Kretschmer, Fiona R. Spuler +2

Predictions on subseasonal-to-seasonal (S2S) timescales--ranging from two weeks to two month--are crucial for early warning systems but remain challenging owing to chaos in the cli…

cs.LG2024

CoSy: Evaluating Textual Explanations of Neurons

Laura Kopf, Philine Lou Bommer, Anna Hedström +3

A crucial aspect of understanding the complex nature of Deep Neural Networks (DNNs) is the ability to explain learned concepts within their latent representations. While methods ex…