A Comparative Study of AI-Generated (GPT-4) and Human-crafted MCQs in Programming Education
arXiv:2312.03173 · doi:10.1145/3636243.3636256
Abstract
There is a constant need for educators to develop and maintain effective up-to-date assessments. While there is a growing body of research in computing education on utilizing large language models (LLMs) in generation and engagement with coding exercises, the use of LLMs for generating programming MCQs has not been extensively explored. We analyzed the capability of GPT-4 to produce multiple-choice questions (MCQs) aligned with specific learning objectives (LOs) from Python programming classes in higher education. Specifically, we developed an LLM-powered (GPT-4) system for generation of MCQs from high-level course context and module-level LOs. We evaluated 651 LLM-generated and 449 human-crafted MCQs aligned to 246 LOs from 6 Python courses. We found that GPT-4 was capable of producing MCQs with clear language, a single correct choice, and high-quality distractors. We also observed that the generated MCQs appeared to be well-aligned with the LOs. Our findings can be leveraged by educators wishing to take advantage of the state-of-the-art generative models to support MCQ authoring efforts.
References in corpus (13)
- Automatic Generation of Programming Exercises and Code Explanations using Large Language Models
- The Robots are Here: Navigating the Generative AI Revolution in Computing Education
- Studying the effect of AI Code Generators on Supporting Novice Learners in Introductory Programming
- Comparing Code Explanations Created by Students and Large Language Models
- Question Answering and Question Generation as Dual Tasks
- Thrilled by Your Progress! Large Language Models (GPT-4) No Longer Struggle to Pass Assessments in Higher Education Programming Courses
- Can Generative Pre-trained Transformers (GPT) Pass Assessments in Higher Education Programming Courses?
- Can GPT-4 Support Analysis of Textual Data in Tasks Requiring Highly Specialized Domain Expertise?
- Patterns of Student Help-Seeking When Using a Large Language Model-Powered Programming Assistant
- Harnessing LLMs in Curricular Design: Using GPT-4 to Support Authoring of Learning Objectives
- Prototyping the use of Large Language Models (LLMs) for adult learning content creation at scale
- Understanding the Role of Temperature in Diverse Question Generation by GPT-4
- Efficient Classification of Student Help Requests in Programming Courses Using Large Language Models
Cited by in corpus (8)
- Automating Autograding: Large Language Models as Test Suite Generators for Introductory Programming
- Understanding the Role of Temperature in Diverse Question Generation by GPT-4
- Generating AI Literacy MCQs: A Multi-Agent LLM Approach
- Evaluating the capability of large language models to personalize science texts for diverse middle-school-age learners
- Synthetic Students: A Comparative Study of Bug Distribution Between Large Language Models and Computing Students
- From Pilots to Practices: A Scoping Review of GenAI-Enabled Personalization in Computer Science Education
- Automating Personalized Parsons Problems with Customized Contexts and Concepts
- Automatic Item Generation for Personality Situational Judgment Tests with Large Language Models