Will Your Next Pair Programming Partner Be Human? An Empirical Evaluation of Generative AI as a Collaborative Teammate in a Semester-Long Classroom Setting
arXiv:2505.08119 · doi:10.1145/3698205.3729544
Abstract
Generative AI (GenAI), especially Large Language Models (LLMs), is rapidly reshaping both programming workflows and computer science education. Many programmers now incorporate GenAI tools into their workflows, including for collaborative coding tasks such as pair programming. While prior research has demonstrated the benefits of traditional pair programming and begun to explore GenAI-assisted coding, the role of LLM-based tools as collaborators in pair programming remains underexamined. In this work, we conducted a mixed-methods study with 39 undergraduate students to examine how GenAI influences collaboration, learning, and performance in pair programming. Specifically, students completed six in-class assignments under three conditions: Traditional Pair Programming (PP), Pair Programming with GenAI (PAI), and Solo Programming with GenAI (SAI). They used both LLM-based inline completion tools (e.g., GitHub Copilot) and LLM-based conversational tools (e.g., ChatGPT). Our results show that students in PAI achieved the highest assignment scores, whereas those in SAI attained the lowest. Additionally, students' attitudes toward LLMs' programming capabilities improved significantly after collaborating with LLM-based tools, and preferences were largely shaped by the perceived usefulness for completing assignments and learning programming skills, as well as the quality of collaboration. Our qualitative findings further reveal that while students appreciated LLM-based tools as valuable pair programming partners, they also identified limitations and had different expectations compared to human teammates. Our study provides one of the first empirical evaluations of GenAI as a pair programming collaborator through a comparison of three conditions (PP, PAI, and SAI). We also discuss the design implications and pedagogical considerations for future GenAI-assisted pair programming approaches.
Accepted by Learning @ Scale 2025
References in corpus (13)
- Automatic Generation of Programming Exercises and Code Explanations using Large Language Models
- Studying the effect of AI Code Generators on Supporting Novice Learners in Introductory Programming
- Using Large Language Models to Enhance Programming Error Messages
- An Empirical Study of the Non-determinism of ChatGPT in Code Generation
- Evaluating the Effectiveness of LLMs in Introductory Computer Science Education: A Semester-Long Field Study
- Desirable Characteristics for AI Teaching Assistants in Programming Education
- Generative AI Assistants in Software Development Education: A vision for integrating Generative AI into educational practice, not instinctively defending against it
- Exploring the Role of AI Assistants in Computer Science Education: Methods, Implications, and Instructor Perspectives
- A Complete Survey on LLM-based AI Chatbots
- Is AI the better programming partner? Human-Human Pair Programming vs. Human-AI pAIr Programming
- Can ChatGPT Play the Role of a Teaching Assistant in an Introductory Programming Course?
- How Novice Programmers Use and Experience ChatGPT when Solving Programming Exercises in an Introductory Course
- Analyzing Chat Protocols of Novice Programmers Solving Introductory Programming Tasks with ChatGPT