2 papers
cs.LG2026
Beyond a Bag of Features: Set-Level Instability in Sparse Autoencoders
Nikolai Bolik, Lennart Stöpler, Artur Andrzejak
Shani et al. (2026) show that LLM representations broadly recover human category boundaries, while failing to reflect fine-grained typicality structure. Their analysis uses cosine…
cs.CL2026
KnowledgeDebugger -- an Exploration Tool for Knowledge Localization and Editing in Transformers
Eric Benz, Lennart Stöpler, Nikolai Bolik +1
Recent research has increasingly focused on understanding how Transformers store and process knowledge, as well as how this knowledge can be edited. Research work in this area is o…