Post-hoc Interpretability for Neural NLP: A Survey
arXiv:2108.04840 · doi:10.1145/3546577
Abstract
Neural networks for NLP are becoming increasingly complex and widespread, and there is a growing concern if these models are responsible to use. Explaining models helps to address the safety and ethical concerns and is essential for accountability. Interpretability serves to provide these explanations in terms that are understandable to humans. Additionally, post-hoc methods provide explanations after a model is learned and are generally model-agnostic. This survey provides a categorization of how recent post-hoc interpretability methods communicate explanations to humans, it discusses each method in-depth, and how they are validated, as the latter is often a common concern.
References in corpus (3)
Cited by in corpus (11)
- Rationalization for Explainable NLP: A Survey
- Inseq: An Interpretability Toolkit for Sequence Generation Models
- ferret: a Framework for Benchmarking Explainers on Transformers
- Tensor networks for interpretable and efficient quantum-inspired machine learning
- Review of Natural Language Processing in Pharmacology
- Beyond Fidelity: Explaining Vulnerability Localization of Learning-based Detectors
- A Comprehensive Survey on Self-Interpretable Neural Networks
- Are Character-level Translations Worth the Wait? Comparing ByT5 and mT5 for Machine Translation
- Interpreting What Typical Fault Signals Look Like via Prototype-matching
- CAT: Interpretable Concept-based Taylor Additive Models
- Segmentation and Smoothing Affect Explanation Quality More Than the Choice of Perturbation-based XAI Method for Image Explanations