A Comprehensive Survey on Self-Interpretable Neural Networks
arXiv:2501.15638 · doi:10.1109/JPROC.2025.3635153
Abstract
Neural networks have achieved remarkable success across various fields. However, the lack of interpretability limits their practical use, particularly in critical decision-making scenarios. Post-hoc interpretability, which provides explanations for pre-trained models, is often at risk of robustness and fidelity. This has inspired a rising interest in self-interpretable neural networks, which inherently reveal the prediction rationale through the model structures. Although there exist surveys on post-hoc interpretability, a comprehensive and systematic survey of self-interpretable neural networks is still missing. To address this gap, we first collect and review existing works on self-interpretable neural networks and provide a structured summary of their methodologies from five key perspectives: attribution-based, function-based, concept-based, prototype-based, and rule-based self-interpretation. We also present concrete, visualized examples of model explanations and discuss their applicability across diverse scenarios, including image, text, graph data, and deep reinforcement learning. Additionally, we summarize existing evaluation metrics for self-interpretability and identify open challenges in this field, offering insights for future research. To support ongoing developments, we present a publicly accessible resource to track advancements in this domain: https://github.com/yangji721/Awesome-Self-Interpretable-Neural-Network.
References in corpus (25)
- A Survey on Large Language Model based Autonomous Agents
- Explaining Deep Neural Networks and Beyond: A Review of Methods and Applications
- Searching for Exotic Particles in High-Energy Physics with Deep Learning
- A Survey on the Explainability of Supervised Machine Learning
- What Do We Want From Explainable Artificial Intelligence (XAI)? -- A Stakeholder Perspective on XAI and a Conceptual Model Guiding Interdisciplinary XAI Research
- Concept Whitening for Interpretable Image Recognition
- Towards Human-centered Explainable AI: A Survey of User Studies for Model Explanations
- A Survey of the State of Explainable AI for Natural Language Processing
- Post-hoc Interpretability for Neural NLP: A Survey
- Explainable Deep Reinforcement Learning: State of the Art and Challenges
- AlphaStock: A Buying-Winners-and-Selling-Losers Investment Strategy using Interpretable Deep Reinforcement Attention Networks
- On the Explainability of Natural Language Processing Deep Models
- TimeSHAP: Explaining Recurrent Models through Sequence Perturbations
- Interpretable and Steerable Sequence Learning via Prototypes
- ProtoPShare: Prototype Sharing for Interpretable Image Classification and Similarity Discovery
- Logic Explained Networks
- Modeling Users' Behavior Sequences with Hierarchical Explainable Network for Cross-domain Fraud Detection
- Entropy-based Logic Explanations of Neural Networks
- Augmenting Interpretable Models with LLMs during Training
- Local Interpretations for Explainable Natural Language Processing: A Survey
- Learning Interpretable Rules for Scalable Data Representation and Classification
- Unveiling Global Interactive Patterns across Graphs: Towards Interpretable Graph Neural Networks
- Extending Logic Explained Networks to Text Classification
- CAT: Interpretable Concept-based Taylor Additive Models
- Towards Robust Metrics for Concept Representation Evaluation