That's Not the Feedback I Need! -- Student Engagement with GenAI Feedback in the Tutor Kai
arXiv:2506.20433 · doi:10.1145/3754508.3754512
Abstract
The potential of Generative AI (GenAI) for generating feedback in computing education has been the subject of numerous studies. However, there is still limited research on how computing students engage with this feedback and to what extent it supports their problem-solving. For this reason, we built a custom web application providing students with Python programming tasks, a code editor, GenAI feedback, and compiler feedback. Via a think-aloud protocol including eye-tracking and a post-interview with 11 undergraduate students, we investigate (1) how much attention the generated feedback received from learners and (2) to what extent the generated feedback is helpful (or not). In addition, students' attention to GenAI feedback is compared with that towards the compiler feedback. We further investigate differences between students with and without prior programming experience. The findings indicate that GenAI feedback generally receives a lot of visual attention, with inexperienced students spending twice as much fixation time. More experienced students requested GenAI less frequently, and could utilize it better to solve the given problem. It was more challenging for inexperienced students to do so, as they could not always comprehend the GenAI feedback. They often relied solely on the GenAI feedback, while compiler feedback was not read. Understanding students' attention and perception toward GenAI feedback is crucial for developing educational tools that support student learning.
Accepted for the UK and Ireland Computing Education Research conference (UKICER 2025)
References in corpus (16)
- CodeAid: Evaluating a Classroom Deployment of an LLM-based Programming Assistant that Balances Student and Educator Needs
- Using Large Language Models to Enhance Programming Error Messages
- Exploring the Responses of Large Language Models to Beginner Programmers' Help Requests
- Evaluating the Effectiveness of LLMs in Introductory Computer Science Education: A Semester-Long Field Study
- The Prompt Report: A Systematic Survey of Prompt Engineering Techniques
- Desirable Characteristics for AI Teaching Assistants in Programming Education
- Feedback-Generation for Programming Exercises With GPT-4
- Exploring How Multiple Levels of GPT-Generated Programming Hints Support or Disappoint Novices
- Evaluating the Application of Large Language Models to Generate Feedback in Programming Education
- Large Language Models in Introductory Programming Education: ChatGPT's Performance and Implications for Assessments
- Not the Silver Bullet: LLM-enhanced Programming Error Messages are Ineffective in Practice
- Leveraging Lecture Content for Improved Feedback: Explorations with GPT-4 and Retrieval Augmented Generation
- Investigating Student Reasoning in Method-Level Code Refactoring: A Think-Aloud Study
- Evaluating Language Models for Generating and Judging Programming Feedback
- Unlimited Practice Opportunities: Automated Generation of Comprehensive, Personalized Programming Tasks
- Prompts First, Finally