Performance and Metacognition Disconnect when Reasoning in Human-AI Interaction
arXiv:2409.16708 · doi:10.1016/j.chb.2025.108779
Abstract
Optimizing human-AI interaction requires users to reflect on their own performance critically. Our paper examines whether people using AI to complete tasks can accurately monitor how well they perform. In Study 1, participants (N = 246) used AI to solve 20 logical problems from the Law School Admission Test. While their task performance improved by three points compared to a norm population, participants overestimated their performance by four points. Interestingly, higher AI literacy was linked to less accurate self-assessment. Participants with more technical knowledge of AI were more confident but less precise in judging their own performance. Using a computational model, we explored individual differences in metacognitive accuracy and found that the Dunning-Kruger effect, usually observed in this task, ceased to exist with AI. Study 2 (N = 452) replicates these findings. We discuss how AI levels metacognitive performance and consider consequences of performance overestimation for interactive AI systems enhancing cognition.
27 pages, 10 figures, 6 tables
References in corpus (7)
- To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-assisted Decision-making
- When combinations of humans and AI are useful: A systematic review and meta-analysis
- The AI Ghostwriter Effect: When Users Do Not Perceive Ownership of AI-Generated Text But Self-Declare as Authors
- The Placebo Effect of Artificial Intelligence in Human-Computer Interaction
- Gender, Age, and Technology Education Influence the Adoption and Appropriation of LLMs
- Multi-Perspective Stance Detection
- People over trust AI-generated medical responses and view them to be as valid as doctors, despite low accuracy