Unmasking Clever Hans Predictors and Assessing What Machines Really Learn
arXiv:1902.10178 · doi:10.1038/s41467-019-08987-4
Abstract
Current learning machines have successfully solved hard application problems, reaching high accuracy and displaying seemingly "intelligent" behavior. Here we apply recent techniques for explaining decisions of state-of-the-art learning machines and analyze various tasks from computer vision and arcade games. This showcases a spectrum of problem-solving behaviors ranging from naive and short-sighted, to well-informed and strategic. We observe that standard performance evaluation metrics can be oblivious to distinguishing these diverse problem solving behaviors. Furthermore, we propose our semi-automated Spectral Relevance Analysis that provides a practically effective way of characterizing and validating the behavior of nonlinear learning machines. This helps to assess whether a learned model indeed delivers reliably for the problem that it was conceived for. Furthermore, our work intends to add a voice of caution to the ongoing excitement about machine intelligence and pledges to evaluate and judge some of these recent successes in a more nuanced manner.
Accepted for publication in Nature Communications
References in corpus (2)
Cited by in corpus (65)
- Machine Learning Force Fields
- Machine learning for molecular simulation
- A Unifying Review of Deep and Shallow Anomaly Detection
- Combining Machine Learning and Computational Chemistry for Predictive Insights Into Chemical Systems
- Towards Explainable Artificial Intelligence
- Explainable Artificial Intelligence (XAI) 2.0: A Manifesto of Open Challenges and Interdisciplinary Research Directions
- Artificial Intelligence and Big Data in Entrepreneurship: A New Era Has Begun
- SpookyNet: Learning Force Fields with Electronic Degrees of Freedom and Nonlocal Effects
- Explanations in Autonomous Driving: A Survey
- The Debate Over Understanding in AI's Large Language Models
- A Comprehensive Taxonomy for Explainable Artificial Intelligence: A Systematic Survey of Surveys on Methods and Concepts
- From Attribution Maps to Human-Understandable Explanations through Concept Relevance Propagation
- Evaluating explainable artificial intelligence methods for multi-label deep learning classification tasks in remote sensing
- Neural Network Potentials for Chemistry: Concepts, Applications and Prospects
- Toward Explainable AI for Regression Models
- Neural Network Attribution Methods for Problems in Geoscience: A Novel Synthetic Benchmark Dataset
- Towards a Collective Agenda on AI for Earth Science Data Analysis
- Implementing local-explainability in Gradient Boosting Trees: Feature Contribution
- Explaining and Interpreting LSTMs
- Don't Push the Button! Exploring Data Leakage Risks in Machine Learning and Transfer Learning
- Responsible and Regulatory Conform Machine Learning for Medicine: A Survey of Challenges and Solutions
- Promises and pitfalls of deep neural networks in neuroimaging-based psychiatric research
- Analysis of a Deep Learning Model for 12-Lead ECG Classification Reveals Learned Features Similar to Diagnostic Criteria
- Machine learning astrophysics from 21 cm lightcones: impact of network architectures and signal contamination
- Survey of XAI in digital pathology
- Machine learning methods for prediction of cancer driver genes: a survey paper
- An XAI framework for robust and transparent data-driven wind turbine power curve models
- Comprehensible Artificial Intelligence on Knowledge Graphs: A survey
- Explaining Deep Learning for ECG Analysis: Building Blocks for Auditing and Knowledge Discovery
- Machine Learning for Observational Cosmology
- The Clever Hans Effect in Unsupervised Learning
- A Survey on Neural Network Interpretability
- White Box Methods for Explanations of Convolutional Neural Networks in Image Classification Tasks
- Discrete and continuous representations and processing in deep learning: Looking forward
- Machine Learning Workflow to Explain Black-box Models for Early Alzheimer's Disease Classification Evaluated for Multiple Datasets
- Improving deep learning with prior knowledge and cognitive models: A survey on enhancing explainability, adversarial robustness and zero-shot learning
- Explainable multiple abnormality classification of chest CT volumes
- Insights Into the Inner Workings of Transformer Models for Protein Function Prediction
- SLISEMAP: Supervised dimensionality reduction through local explanations
- Disentangled Explanations of Neural Network Predictions by Finding Relevant Subspaces
- OnRAMP for Regulating AI in Medical Products
- Machine Learning in Biomechanics: Key Applications and Limitations in Walking, Running, and Sports Movements
- Expressive Explanations of DNNs by Combining Concept Analysis with ILP
- Next2You: Robust Copresence Detection Based on Channel State Information
- The Importance of Being Interpretable: Toward An Understandable Machine Learning Encoder for Galaxy Cluster Cosmology
- PredDiff: Explanations and Interactions from Conditional Expectations
- MAPS-X: Explainable Multi-Robot Motion Planning via Segmentation
- Reliability Scores from Saliency Map Clusters for Improved Image-based Harvest-Readiness Prediction in Cauliflower
- Insightful analysis of historical sources at scales beyond human capabilities using unsupervised Machine Learning and XAI
- Human-Centered Evaluation of XAI Methods
- i-Align: an interpretable knowledge graph alignment model
- A Conceptual Framework for Establishing Trust in Real World Intelligent Systems
- Study on the Helpfulness of Explainable Artificial Intelligence
- Preemptively Pruning Clever-Hans Strategies in Deep Neural Networks
- Computationally efficient neural network classifiers for next generation closed loop neuromodulation therapy -- a case study in epilepsy
- On The Coherence of Quantitative Evaluation of Visual Explanations
- Exploring complex pattern formation with convolutional neural networks
- SE3D: A Framework For Saliency Method Evaluation In 3D Imaging
- What Do Large Language Models Know? Tacit Knowledge as a Potential Causal-Explanatory Structure
- Local Concept Embeddings for Analysis of Concept Distributions in Vision DNN Feature Spaces
- Space-scale Exploration of the Poor Reliability of Deep Learning Models: the Case of the Remote Sensing of Rooftop Photovoltaic Systems
- Towards Desiderata-Driven Design of Visual Counterfactual Explainers
- P-TAME: Explain Any Image Classifier with Trained Perturbations
- Multi-Scale Grouped Prototypes for Interpretable Semantic Segmentation
- How can we trust opaque systems? Criteria for robust explanations in XAI