2 papers
cs.LG2026
Do Sparse Autoencoders Learn Meaningful Concept Hierarchies?
Nils Grandien, David Steinmann, Felix Friedrich +1
Sparse autoencoders (SAEs) have become an important tool for unsupervised concept discovery in large models. To make the resulting feature spaces more interpretable and manageable,…
cs.AI2025
Interpretable end-to-end Neurosymbolic Reinforcement Learning agents
Nils Grandien, Quentin Delfosse, Kristian Kersting
Deep reinforcement learning (RL) agents rely on shortcut learning, preventing them from generalizing to slightly different environments. To address this problem, symbolic method, t…