Investigating and Designing for Trust in AI-powered Code Generation Tools
arXiv:2305.11248 · doi:10.1145/3630106.3658984
Abstract
As AI-powered code generation tools such as GitHub Copilot become popular, it is crucial to understand software developers' trust in AI tools -- a key factor for tool adoption and responsible usage. However, we know little about how developers build trust with AI, nor do we understand how to design the interface of generative AI systems to facilitate their appropriate levels of trust. In this paper, we describe findings from a two-stage qualitative investigation. We first interviewed 17 developers to contextualize their notions of trust and understand their challenges in building appropriate trust in AI code generation tools. We surfaced three main challenges -- including building appropriate expectations, configuring AI tools, and validating AI suggestions. To address these challenges, we conducted a design probe study in the second stage to explore design concepts that support developers' trust-building process by 1) communicating AI performance to help users set proper expectations, 2) allowing users to configure AI by setting and adjusting preferences, and 3) offering indicators of model mechanism to support evaluation of AI suggestions. We gathered developers' feedback on how these design concepts can help them build appropriate trust in AI-powered code generation tools, as well as potential risks in design. These findings inform our proposed design recommendations on how to design for trust in AI-powered code generation tools.
accepted to FAccT 2024
References in corpus (18)
- To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-assisted Decision-making
- Questioning the AI: Informing Design Practices for Explainable AI User Experiences
- Effect of Confidence and Explanation on Accuracy and Trust Calibration in AI-Assisted Decision Making
- Opportunities and Challenges in Explainable Artificial Intelligence (XAI): A Survey
- CoAuthor: Designing a Human-AI Collaborative Writing Dataset for Exploring Language Model Capabilities
- Proxy Tasks and Subjective Measures Can Be Misleading in Evaluating Explainable AI Systems
- Designing for Responsible Trust in AI Systems: A Communication Perspective
- Trust in AutoML: Exploring Information Needs for Establishing Trust in Automated Machine Learning Systems
- Friend, Collaborator, Student, Manager: How Design of an AI-Driven Game Level Editor Affects Creators
- Perfection Not Required? Human-AI Partnerships in Code Translation
- Better Together? An Evaluation of AI-Supported Code Translation
- AI-driven Development Is Here: Should You Worry?
- What is it like to program with artificial intelligence?
- GitHub Copilot AI pair programmer: Asset or Liability?
- Humans, AI, and Context: Understanding End-Users' Trust in a Real-World Computer Vision Application
- Grounded Copilot: How Programmers Interact with Code-Generating Models
- Reading Between the Lines: Modeling User Behavior and Costs in AI-Assisted Programming
- Quality Estimation & Interpretability for Code Translation
Cited by in corpus (5)
- The Impact of Generative AI Coding Assistants on Developers Who Are Visually Impaired
- Human-AI Experience in Integrated Development Environments: A Systematic Literature Review
- A Transformer-Based Approach for Smart Invocation of Automatic Code Completion
- AI Trust Reshaping Administrative Burdens: Understanding Trust-Burden Dynamics in LLM-Assisted Benefits Systems
- Towards More Trustworthy and Interpretable LLMs for Code through Syntax-Grounded Explanations