How Beginning Programmers and Code LLMs (Mis)read Each Other
arXiv:2401.15232 · doi:10.1145/3613904.3642706
Abstract
Generative AI models, specifically large language models (LLMs), have made strides towards the long-standing goal of text-to-code generation. This progress has invited numerous studies of user interaction. However, less is known about the struggles and strategies of non-experts, for whom each step of the text-to-code problem presents challenges: describing their intent in natural language, evaluating the correctness of generated code, and editing prompts when the generated code is incorrect. This paper presents a large-scale controlled study of how 120 beginning coders across three academic institutions approach writing and editing prompts. A novel experimental design allows us to target specific steps in the text-to-code process and reveals that beginners struggle with writing and editing prompts, even for problems at their skill level and when correctness is automatically determined. Our mixed-methods evaluation provides insight into student processes and perceptions with key implications for non-expert Code LLM use within and outside of education.
Published in CHI 2024
References in corpus (11)
- Studying the effect of AI Code Generators on Supporting Novice Learners in Introductory Programming
- The Programmer's Assistant: Conversational Interaction with a Large Language Model for Software Development
- CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis
- Using Large Language Models to Enhance Programming Error Messages
- "It's Weird That it Knows What I Want": Usability and Interactions with Copilot for Novice Programmers
- Comparing Code Explanations Created by Students and Large Language Models
- A Case for Humans-in-the-Loop: Decisions in the Presence of Erroneous Algorithmic Scores
- "What It Wants Me To Say": Bridging the Abstraction Gap Between End-User Programmers and Code-Generating Large Language Models
- An Empirical Evaluation of Using Large Language Models for Automated Unit Test Generation
- GitHub Copilot AI pair programmer: Asset or Liability?
- Data Race Detection Using Large Language Models
Cited by in corpus (9)
- What Should We Engineer in Prompts? Training Humans in Requirement-Driven LLM Use
- Controlling AI Agent Participation in Group Conversations: A Human-Centered Approach
- The Impact of Generative AI Coding Assistants on Developers Who Are Visually Impaired
- Not the Silver Bullet: LLM-enhanced Programming Error Messages are Ineffective in Practice
- How Scientists Use Large Language Models to Program
- InstructPipe: Generating Visual Blocks Pipelines with Human Instructions and LLMs
- Gensors: Authoring Personalized Visual Sensors with Multimodal Foundation Models and Reasoning
- Howzat? Appealing to Expert Judgement for Evaluating Human and AI Next-Step Hints for Novice Programmers
- Examining the Usage of Generative AI Models in Student Learning Activities for Software Programming