15 citations · 15 across the 15 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Capturing Polysemanticity with PRISM: A Multi-Concept Feature Description Framework
Laura Kopf, Nils Feldhus, Kirill Bykov +4
Automated interpretability research aims to identify concepts encoded in neural network features to enhance human understanding of model behavior. Within the context of large langu…
cs.LG2021
Efficient Explanations from Empirical Explainers
Robert Schwarzenberg, Nils Feldhus, Sebastian Möller
Amid a discussion about Green AI in which we see explainability neglected, we explore the possibility to efficiently approximate computationally expensive explainers. To this end,…