People Attribute Purpose to Autonomous Vehicles When Explaining Their Behavior: Insights from Cognitive Science for Explainable AI
arXiv:2403.08828 · doi:10.1145/3706598.3713509
Abstract
It is often argued that effective human-centered explainable artificial intelligence (XAI) should resemble human reasoning. However, empirical investigations of how concepts from cognitive science can aid the design of XAI are lacking. Based on insights from cognitive science, we propose a framework of explanatory modes to analyze how people frame explanations, whether mechanistic, teleological, or counterfactual. Using the complex safety-critical domain of autonomous driving, we conduct an experiment consisting of two studies on (i) how people explain the behavior of a vehicle in 14 unique scenarios (N1=54) and (ii) how they perceive these explanations (N2=382), curating the novel Human Explanations for Autonomous Driving Decisions (HEADD) dataset. Our main finding is that participants deem teleological explanations significantly better quality than counterfactual ones, with perceived teleology being the best predictor of perceived quality. Based on our results, we argue that explanatory modes are an important axis of analysis when designing and evaluating XAI and highlight the need for a principled and empirically grounded understanding of the cognitive mechanisms of explanation. The HEADD dataset and our code are available at: https://datashare.ed.ac.uk/handle/10283/8930.
CHI 2025
References in corpus (7)
- Proxy Tasks and Subjective Measures Can Be Misleading in Evaluating Explainable AI Systems
- "Help Me Help the AI": Understanding How Explainability Can Support Human-AI Interaction
- Explainable AI for Safe and Trustworthy Autonomous Driving: A Systematic Review
- Impossibility Theorems for Feature Attribution
- Bridging the Transparency Gap: What Can Explainable AI Learn From the AI Act?
- Keep Your Friends Close and Your Counterfactuals Closer: Improved Learning From Closest Rather Than Plausible Counterfactual Explanations in an Abstract Setting
- Integrating Policy Summaries with Reward Decomposition for Explaining Reinforcement Learning Agents