Effect of Confidence and Explanation on Accuracy and Trust Calibration in AI-Assisted Decision Making
arXiv:2001.02114 · doi:10.1145/3351095.3372852
Abstract
Today, AI is being increasingly used to help human experts make decisions in high-stakes scenarios. In these scenarios, full automation is often undesirable, not only due to the significance of the outcome, but also because human experts can draw on their domain knowledge complementary to the model's to ensure task success. We refer to these scenarios as AI-assisted decision making, where the individual strengths of the human and the AI come together to optimize the joint decision outcome. A key to their success is to appropriately \textit{calibrate} human trust in the AI on a case-by-case basis; knowing when to trust or distrust the AI allows the human expert to appropriately apply their knowledge, improving decision outcomes in cases where the model is likely to perform poorly. This research conducts a case study of AI-assisted decision making in which humans and AI have comparable performance alone, and explores whether features that reveal case-specific model information can calibrate trust and improve the joint performance of the human and AI. Specifically, we study the effect of showing confidence score and local explanation for a particular prediction. Through two human experiments, we show that confidence score can help calibrate people's trust in an AI model, but trust calibration alone is not sufficient to improve AI-assisted decision making, which may also depend on whether the human can bring in enough unique knowledge to complement the AI's errors. We also highlight the problems in using local explanation for AI-assisted decision making scenarios and invite the research community to explore new approaches to explainability for calibrating human trust in AI.
References in corpus (1)
Cited by in corpus (66)
- To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-assisted Decision-making
- Expanding Explainability: Towards Social Transparency in AI systems
- When combinations of humans and AI are useful: A systematic review and meta-analysis
- Towards Human-centered Explainable AI: A Survey of User Studies for Model Explanations
- Explainable Artificial Intelligence: A Survey of Needs, Techniques, Applications, and Future Direction
- "Help Me Help the AI": Understanding How Explainability Can Support Human-AI Interaction
- Appropriate Reliance on AI Advice: Conceptualization and the Effect of Explanations
- Assessing the communication gap between AI models and healthcare professionals: explainability, utility and trust in AI-driven clinical decision-making
- Designing for Responsible Trust in AI Systems: A Communication Perspective
- "I'm Not Sure, But...": Examining the Impact of Large Language Models' Uncertainty Expression on User Reliance and Trust
- Trust in AutoML: Exploring Information Needs for Establishing Trust in Automated Machine Learning Systems
- Beyond Expertise and Roles: A Framework to Characterize the Stakeholders of Interpretable Machine Learning and their Needs
- Interactive Model Cards: A Human-Centered Approach to Model Documentation
- Perfection Not Required? Human-AI Partnerships in Code Translation
- Understanding the Effect of Out-of-distribution Examples and Interactive Explanations on Human-AI Decision Making
- A Meta-Analysis of the Utility of Explainable Artificial Intelligence in Human-AI Decision-Making
- Human-AI Collaboration: The Effect of AI Delegation on Human Task Performance and Task Satisfaction
- Rethinking Human-AI Collaboration in Complex Medical Decision Making: A Case Study in Sepsis Diagnosis
- Improving Human-AI Collaboration With Descriptions of AI Behavior
- Charting the Sociotechnical Gap in Explainable AI: A Framework to Address the Gap in XAI
- Investigating and Designing for Trust in AI-powered Code Generation Tools
- On the Quest for Effectiveness in Human Oversight: Interdisciplinary Perspectives
- XAIR: A Framework of Explainable AI in Augmented Reality
- Explanations, Fairness, and Appropriate Reliance in Human-AI Decision-Making
- Twenty-Four Years of Empirical Research on Trust in AI: A Bibliometric Review of Trends, Overlooked Issues, and Future Directions
- AutoAIViz: Opening the Blackbox of Automated Artificial Intelligence with Conditional Parallel Coordinates
- Fostering Appropriate Reliance on Large Language Models: The Role of Explanations, Sources, and Inconsistencies
- Fact-checking information from large language models can decrease headline discernment
- A Critical Survey on Fairness Benefits of Explainable AI
- Is Conversational XAI All You Need? Human-AI Decision Making With a Conversational XAI Assistant
- Competent but Rigid: Identifying the Gap in Empowering AI to Participate Equally in Group Decision-Making
- Conversate: Supporting Reflective Learning in Interview Practice Through Interactive Simulation and Dialogic Feedback
- Seamful XAI: Operationalizing Seamful Design in Explainable AI
- Accuracy-Time Tradeoffs in AI-Assisted Decision Making under Time Pressure
- One vs. Many: Comprehending Accurate Information from Multiple Erroneous and Inconsistent AI Generations
- Misfitting With AI: How Blind People Verify and Contest AI Errors
- As Confidence Aligns: Exploring the Effect of AI Confidence on Human Self-confidence in Human-AI Decision Making
- AI, Help Me Think$\unicode{x2014}$but for Myself: Assisting People in Complex Decision-Making by Providing Different Kinds of Cognitive Support
- ASHABot: An LLM-Powered Chatbot to Support the Informational Needs of Community Health Workers
- Plan-Then-Execute: An Empirical Study of User Trust and Team Performance When Using LLM Agents As A Daily Assistant
- Using ChatGPT in HCI Research -- A Trioethnography
- Comparing Zealous and Restrained AI Recommendations in a Real-World Human-AI Collaboration Task
- Unraveling the Dilemma of AI Errors: Exploring the Effectiveness of Human and Machine Explanations for Large Language Models
- Bridging Generations using AI-Supported Co-Creative Activities
- Explainable AI Reloaded: Challenging the XAI Status Quo in the Era of Large Language Models
- The AI-DEC: A Card-based Design Method for User-centered AI Explanations
- Beyond Recommendations: From Backward to Forward AI Support of Pilots' Decision-Making Process
- The Impact of Explanations on Fairness in Human-AI Decision-Making: Protected vs Proxy Features
- Media Bias Detector: Designing and Implementing a Tool for Real-Time Selection and Framing Bias Analysis in News Coverage
- Exploring the Effects of Chatbot Anthropomorphism and Human Empathy on Human Prosocial Behavior Toward Chatbots
- A Comparative User Study of Human Predictions in Algorithm-Supported Recidivism Risk Assessment
- (Beyond) Reasonable Doubt: Challenges that Public Defenders Face in Scrutinizing AI in Court
- PADTHAI-MM: Principles-based Approach for Designing Trustworthy, Human-centered AI using MAST Methodology
- Design Considerations for Human Oversight of AI: Insights from Co-Design Workshops and Work Design Theory
- A Matter of Perspective(s): Contrasting Human and LLM Argumentation in Subjective Decision-Making on Subtle Sexism
- Towards Uncertainty Aware Task Delegation and Human-AI Collaborative Decision-Making
- Give Me a Choice: The Consequences of Restricting Choices Through AI-Support for Perceived Autonomy, Motivational Variables, and Decision Performance
- Towards Feature Engineering with Human and AI's Knowledge: Understanding Data Science Practitioners' Perceptions in Human&AI-Assisted Feature Engineering Design
- "Even explanations will not help in trusting [this] fundamentally biased system": A Predictive Policing Case-Study
- AInsight: Augmenting Expert Decision-Making with On-the-Fly Insights Grounded in Historical Data
- ContextualSHAP : Enhancing SHAP Explanations Through Contextual Language Generation
- From Model Performance to Claim: How a Change of Focus in Machine Learning Replicability Can Help Bridge the Responsibility Gap
- The Contribution of XAI for the Safe Development and Certification of AI: An Expert-Based Analysis
- Can Offline Metrics Measure Explanation Goals? A Comparative Survey Analysis of Offline Explanation Metrics in Recommender Systems
- Learning from AVA: Early Lessons from a Curated and Trustworthy Generative AI for Policy and Development Research
- Human-Centered Explainable AI for Security Enhancement: A Deep Intrusion Detection Framework