The Clever Hans Effect in Unsupervised Learning
arXiv:2408.08041 · doi:10.1038/s42256-025-01000-2
Abstract
Unsupervised learning has become an essential building block of AI systems. The representations it produces, e.g. in foundation models, are critical to a wide variety of downstream applications. It is therefore important to carefully examine unsupervised models to ensure not only that they produce accurate predictions, but also that these predictions are not "right for the wrong reasons", the so-called Clever Hans (CH) effect. Using specially developed Explainable AI techniques, we show for the first time that CH effects are widespread in unsupervised learning. Our empirical findings are enriched by theoretical insights, which interestingly point to inductive biases in the unsupervised learning machine as a primary source of CH effects. Overall, our work sheds light on unexplored risks associated with practical applications of unsupervised learning and suggests ways to make unsupervised learning more robust.
12 pages + supplement
References in corpus (19)
- Learning Transferable Visual Models From Natural Language Supervision
- ChestX-ray8: Hospital-scale Chest X-ray Database and Benchmarks on Weakly-Supervised Classification and Localization of Common Thorax Diseases
- Methods for Interpreting and Understanding Deep Neural Networks
- Deep Learning for Anomaly Detection: A Review
- On the Opportunities and Risks of Foundation Models
- Shortcut Learning in Deep Neural Networks
- Explaining Deep Neural Networks and Beyond: A Review of Methods and Applications
- Unmasking Clever Hans Predictors and Assessing What Machines Really Learn
- A Unifying Review of Deep and Shallow Anomaly Detection
- Higher-Order Explanations of Graph Neural Networks via Relevant Walks
- Towards Explaining Anomalies: A Deep Taylor Decomposition of One-Class Models
- Explaining and Interpreting LSTMs
- Building and Interpreting Deep Similarity Models
- Frequency Bias in Neural Networks for Input of Non-Uniform Density
- XAI for Transformers: Better Explanations through Conservative Propagation
- RudolfV: A Foundation Model by Pathologists for Pathologists
- Disentangled Explanations of Neural Network Predictions by Finding Relevant Subspaces
- Insightful analysis of historical sources at scales beyond human capabilities using unsupervised Machine Learning and XAI
- Preemptively Pruning Clever-Hans Strategies in Deep Neural Networks