Automated Correction for Syntax Errors in Programming Assignments using Recurrent Neural Networks
arXiv:1603.06129
Abstract
We present a method for automatically generating repair feedback for syntax errors for introductory programming problems. Syntax errors constitute one of the largest classes of errors (34%) in our dataset of student submissions obtained from a MOOC course on edX. The previous techniques for generating automated feed- back on programming assignments have focused on functional correctness and style considerations of student programs. These techniques analyze the program AST of the program and then perform some dynamic and symbolic analyses to compute repair feedback. Unfortunately, it is not possible to generate ASTs for student pro- grams with syntax errors and therefore the previous feedback techniques are not applicable in repairing syntax errors. We present a technique for providing feedback on syntax errors that uses Recurrent neural networks (RNNs) to model syntactically valid token sequences. Our approach is inspired from the recent work on learning language models from Big Code (large code corpus). For a given programming assignment, we first learn an RNN to model all valid token sequences using the set of syntactically correct student submissions. Then, for a student submission with syntax errors, we query the learnt RNN model with the prefix to- ken sequence to predict token sequences that can fix the error by either replacing or inserting the predicted token sequence at the error location. We evaluate our technique on over 14, 000 student submissions with syntax errors. Our technique can completely re- pair 31.69% (4501/14203) of submissions with syntax errors and in addition partially correct 6.39% (908/14203) of the submissions.
References in corpus (2)
Cited by in corpus (24)
- Sorting and Transforming Program Repair Ingredients via Deep Learning Code Similarities
- Big Code != Big Vocabulary: Open-Vocabulary Models for Source Code
- Not all bytes are equal: Neural byte sieve for fuzzing
- Dynamic Neural Program Embedding for Program Repair
- Maybe Deep Neural Networks are the Best Choice for Modeling Source Code
- Context2Name: A Deep Learning-Based Approach to Infer Natural Variable Names from Usage Contexts
- Learn&Fuzz: Machine Learning for Input Fuzzing
- A Survey of Automated Programming Hint Generation -- The HINTS Framework
- Learning to Generate Corrective Patches using Neural Machine Translation
- Semantic Code Repair using Neuro-Symbolic Transformation Networks
- Neural Bug Finding: A Study of Opportunities and Challenges
- DeepBugs: A Learning Approach to Name-based Bug Detection
- End-to-End Prediction of Buffer Overruns from Raw Source Code via Neural Memory Networks
- Neutaint: Efficient Dynamic Taint Analysis with Neural Networks
- Towards Proof Synthesis Guided by Neural Machine Translation for Intuitionistic Propositional Logic
- Tailored Mutants Fit Bugs Better
- Variable Name Recovery in Decompiled Binary Code using Constrained Masked Language Modeling
- Language Modelling for Source Code with Transformer-XL
- Fault Localization with Code Coverage Representation Learning
- Play to Grade: Testing Coding Games as Classifying Markov Decision Process
- Embedding Code Contexts for Cryptographic API Suggestion:New Methodologies and Comparisons
- Automatic Repair and Type Binding of Undeclared Variables using Neural Networks
- Assessing the Effectiveness of Syntactic Structure to Learn Code Edit Representations
- You Cannot Fix What You Cannot Find! An Investigation of Fault Localization Bias in Benchmarking Automated Program Repair Systems