collaborators

6 papers

cs.CL2026

A Systematic Comparison between Extractive Self-Explanations and Human Rationales in Text Classification

Stephanie Brandl, Oliver Eberle

Instruction-tuned LLMs are able to provide \textit{an} explanation about their output to users by generating self-explanations, without requiring the application of complex interpr…

cs.CV2026

Beyond Attention Heatmaps: How to Get Better Explanations for Multiple Instance Learning Models in Histopathology

Mina Jamshidi Idaji, Julius Hense, Tom Neuhäuser +12

Multiple instance learning (MIL) has enabled substantial progress in computational histopathology, where a large amount of patches from gigapixel whole slide images are aggregated…

cs.LG2025

Capturing Polysemanticity with PRISM: A Multi-Concept Feature Description Framework

Laura Kopf, Nils Feldhus, Kirill Bykov +4

Automated interpretability research aims to identify concepts encoded in neural network features to enhance human understanding of model behavior. Within the context of large langu…

cs.LG2025

RelP: Faithful and Efficient Circuit Discovery in Language Models via Relevance Patching

Farnoush Rezaei Jafari, Oliver Eberle, Ashkan Khakzar +1

Activation patching is a standard method in mechanistic interpretability for localizing the components of a model responsible for specific behaviors, but it is computationally expe…

cs.CL2025

Trick or Neat: Adversarial Ambiguity and Language Model Evaluation

Antonia Karamolegkou, Oliver Eberle, Phillip Rust +2

Detecting ambiguity is important for language understanding, including uncertainty estimation, humour detection, and processing garden path sentences. We assess language models' se…

cs.LG2025

Cat, Rat, Meow: On the Alignment of Language Model and Human Term-Similarity Judgments

Lorenz Linhardt, Tom Neuhäuser, Lenka Tětková +1

Small and mid-sized generative language models have gained increasing attention. Their size and availability make them amenable to being analyzed at a behavioral as well as a repre…