Abstraction Alignment: Comparing Model-Learned and Human-Encoded Conceptual Relationships
arXiv:2407.12543 · doi:10.1145/3706598.3713406
Abstract
While interpretability methods identify a model's learned concepts, they overlook the relationships between concepts that make up its abstractions and inform its ability to generalize to new data. To assess whether models' have learned human-aligned abstractions, we introduce abstraction alignment, a methodology to compare model behavior against formal human knowledge. Abstraction alignment externalizes domain-specific human knowledge as an abstraction graph, a set of pertinent concepts spanning levels of abstraction. Using the abstraction graph as a ground truth, abstraction alignment measures the alignment of a model's behavior by determining how much of its uncertainty is accounted for by the human abstractions. By aggregating abstraction alignment across entire datasets, users can test alignment hypotheses, such as which human concepts the model has learned and where misalignments recur. In evaluations with experts, abstraction alignment differentiates seemingly similar errors, improves the verbosity of existing model-quality metrics, and uncovers improvements to current human abstractions.
20 pages, 7 figures, published in CHI 2025
References in corpus (35)
- Distilling the Knowledge in a Neural Network
- Towards A Rigorous Science of Interpretable Machine Learning
- A Survey on Knowledge Graphs: Representation, Acquisition and Applications
- A Review of Uncertainty Quantification in Deep Learning: Techniques, Applications and Challenges
- Hierarchical Text-Conditional Image Generation with CLIP Latents
- Aleatoric and Epistemic Uncertainty in Machine Learning: An Introduction to Concepts and Methods
- Explaining Machine Learning Classifiers through Diverse Counterfactual Explanations
- The What-If Tool: Interactive Probing of Machine Learning Models
- Power to the People? Opportunities and Challenges for Participatory AI
- Beyond Expertise and Roles: A Framework to Characterize the Stakeholders of Interpretable Machine Learning and their Needs
- Data Feminism for AI
- Gender Bias in Word Embeddings: A Comprehensive Analysis of Frequency, Syntax, and Semantics
- Understanding Practices, Challenges, and Opportunities for User-Engaged Algorithm Auditing in Industry Practice
- Automated Medical Coding on MIMIC-III and MIMIC-IV: A Critical Review and Replicability Study
- Zeno: An Interactive Framework for Behavioral Evaluation of Machine Learning
- Symphony: Composing Interactive Interfaces for Machine Learning
- Why Does ChatGPT Fall Short in Providing Truthful Answers?
- EXMOS: Explanatory Model Steering Through Multifaceted Explanations and Data Configurations
- Getting aligned on representational alignment
- Interactive AI Alignment: Specification, Process, and Evaluation Alignment
- Angler: Helping Machine Translation Practitioners Prioritize Model Improvements
- From Fitting Participation to Forging Relationships: The Art of Participatory ML
- Inspecting and Editing Knowledge Representations in Language Models
- Saliency Cards: A Framework to Characterize and Compare Saliency Methods
- Compress and Compare: Interactively Evaluating Efficiency and Behavior Across ML Model Compression Experiments
- Concept Alignment
- Bridging the Human-AI Knowledge Gap: Concept Discovery and Transfer in AlphaZero
- Talaria: Interactively Optimizing Machine Learning Models for Efficient Inference
- No Fair Lunch: A Causal Perspective on Dataset Bias in Machine Learning for Medical Imaging
- Aligning Machine and Human Visual Representations across Abstraction Levels
- Lessons Learned from EXMOS User Studies: A Technical Report Summarizing Key Takeaways from User Studies Conducted to Evaluate The EXMOS Platform
- DiffusionWorldViewer: Exposing and Broadening the Worldview Reflected by Generative Text-to-Image Models
- Relational Composition in Neural Networks: A Survey and Call to Action
- eSPARQL: Representing and Reconciling Agnostic and Atheistic Beliefs in RDF-star Knowledge Graphs
- Dimensions of Disagreement: Unpacking Divergence and Misalignment in Cognitive Science and Artificial Intelligence