Automated Grading and Feedback Tools for Programming Education: A Systematic Review
arXiv:2306.11722 · doi:10.1145/3636515
Abstract
We conducted a systematic literature review on automated grading and feedback tools for programming education. We analysed 121 research papers from 2017 to 2021 inclusive and categorised them based on skills assessed, approach, language paradigm, degree of automation and evaluation techniques. Most papers assess the correctness of assignments in object-oriented languages. Typically, these tools use a dynamic technique, primarily unit testing, to provide grades and feedback to the students or static analysis techniques to compare a submission with a reference solution or with a set of correct student submissions. However, these techniques' feedback is often limited to whether the unit tests have passed or failed, the expected and actual output, or how they differ from the reference solution. Furthermore, few tools assess the maintainability, readability or documentation of the source code, with most using static analysis techniques, such as code quality metrics, in conjunction with grading correctness. Additionally, we found that most tools offered fully automated assessment to allow for near-instantaneous feedback and multiple resubmissions, which can increase student satisfaction and provide them with more opportunities to succeed. In terms of techniques used to evaluate the tools' performance, most papers primarily use student surveys or compare the automatic assessment tools to grades or feedback provided by human graders. However, because the evaluation dataset is frequently unavailable, it is more difficult to reproduce results and compare tools to a collection of common assignments.
Accepted version of the manuscript
References in corpus (7)
- Can GPT-4 Support Analysis of Textual Data in Tasks Requiring Highly Specialized Domain Expertise?
- Learning Program Embeddings to Propagate Feedback on Student Code
- Code Quality Evaluation Methodology Using The ISO/IEC 9126 Standard
- Effects of Human vs. Automatic Feedback on Students' Understanding of AI Concepts and Programming Style
- Computing with CodeRunner at Coventry University: Automated summative assessment of Python and C++ code
- A Comparison of Inquiry-Based Conceptual Feedback vs. Traditional Detailed Feedback Mechanisms in Software Testing Education: An Empirical Investigation
- Automatic Assessment of the Design Quality of Python Programs with Personalized Feedback