Automated Classification of Overfitting Patches with Statically Extracted Code Features
arXiv:1910.12057 · doi:10.1109/tse.2021.3071750
Abstract
Automatic program repair (APR) aims to reduce the cost of manually fixing software defects. However, APR suffers from generating a multitude of overfitting patches, those patches that fail to correctly repair the defect beyond making the tests pass. This paper presents a novel overfitting patch detection system called ODS to assess the correctness of APR patches. ODS first statically compares a patched program and a buggy program in order to extract code features at the abstract syntax tree (AST) level. Then, ODS uses supervised learning with the captured code features and patch correctness labels to automatically learn a probabilistic model. The learned ODS model can then finally be applied to classify new and unseen program repair patches. We conduct a large-scale experiment to evaluate the effectiveness of ODS on patch correctness classification based on 10,302 patches from Defects4J, Bugs.jar and Bears benchmarks. The empirical evaluation shows that ODS is able to correctly classify 71.9% of program repair patches from 26 projects, which improves the state-of-the-art. ODS is applicable in practice and can be employed as a post-processing procedure to classify the patches generated by different APR systems.
References in corpus (5)
- Bears: An Extensible Java Bug Benchmark for Automatic Program Repair Studies
- Automated Patch Assessment for Program Repair at Scale
- Learning the Relation between Code Features and Code Transforms with Structured Prediction
- Empirical Review of Java Program Repair Tools: A Large-Scale Experiment on 2,141 Bugs and 23,551 Repair Attempts
- Validation of Automatically Generated Patches: An Appetizer
Cited by in corpus (12)
- Neural Program Repair with Execution-based Backpropagation
- Deep Learning-based Software Engineering: Progress, Challenges, and Opportunities
- SelfAPR: Self-supervised Program Repair with Test Execution Diagnostics
- Invalidator: Automated Patch Correctness Assessment via Semantic and Syntactic Reasoning
- A Software-Repair Robot based on Continual Learning
- ObjSim: Lightweight Automatic Patch Prioritization via Object Similarity
- Is this Change the Answer to that Problem? Correlating Descriptions of Bug and Code Changes for Evaluating Patch Correctness
- Evaluating Representation Learning of Code Changes for Predicting Patch Correctness in Program Repair
- Beep: Fine-grained Fix Localization by Learning to Predict Buggy Code Elements
- Estimating the Potential of Program Repair Search Spaces with Commit Analysis
- RePaCA: Leveraging Reasoning Large Language Models for Static Automated Patch Correctness Assessment
- Show Me Why It's Correct: Saving 1/3 of Debugging Time in Program Repair with Interactive Runtime Comparison