Synthetic Students: A Comparative Study of Bug Distribution Between Large Language Models and Computing Students
arXiv:2410.09193 · doi:10.1145/3649165.3690100
Abstract
Large language models (LLMs) present an exciting opportunity for generating synthetic classroom data. Such data could include code containing a typical distribution of errors, simulated student behaviour to address the cold start problem when developing education tools, and synthetic user data when access to authentic data is restricted due to privacy reasons. In this research paper, we conduct a comparative study examining the distribution of bugs generated by LLMs in contrast to those produced by computing students. Leveraging data from two previous large-scale analyses of student-generated bugs, we investigate whether LLMs can be coaxed to exhibit bug patterns that are similar to authentic student bugs when prompted to inject errors into code. The results suggest that unguided, LLMs do not generate plausible error distributions, and many of the generated errors are unlikely to be generated by real students. However, with guidance including descriptions of common errors and typical frequencies, LLMs can be shepherded to generate realistic distributions of errors in synthetic code.
References in corpus (5)
- Automatic Generation of Programming Exercises and Code Explanations using Large Language Models
- Using Large Language Models to Enhance Programming Error Messages
- Comparing Code Explanations Created by Students and Large Language Models
- A Comparative Study of AI-Generated (GPT-4) and Human-crafted MCQs in Programming Education
- Empirical Evaluation of Deep Learning Models for Knowledge Tracing: Of Hyperparameters and Metrics on Performance and Replicability