Automated Patch Assessment for Program Repair at Scale
arXiv:1909.13694 · doi:10.1007/s10664-020-09920-w
Abstract
In this paper, we do automatic correctness assessment for patches generated by program repair systems. We consider the human-written patch as ground truth oracle and randomly generate tests based on it, a technique proposed by Shamshiri et al., called Random testing with Ground Truth (RGT) in this paper. We build a curated dataset of 638 patches for Defects4J generated by 14 state-of-the-art repair systems, we evaluate automated patch assessment on this dataset. The results of this study are novel and significant: First, we improve the state of the art performance of automatic patch assessment with RGT by 190% by improving the oracle; Second, we show that RGT is reliable enough to help scientists to do overfitting analysis when they evaluate program repair systems; Third, we improve the external validity of the program repair knowledge with the largest study ever.
References in corpus (2)
Cited by in corpus (15)
- Neural Program Repair with Execution-based Backpropagation
- SelfAPR: Self-supervised Program Repair with Test Execution Diagnostics
- Automated Classification of Overfitting Patches with Statically Extracted Code Features
- Invalidator: Automated Patch Correctness Assessment via Semantic and Syntactic Reasoning
- Sorald: Automatic Patch Suggestions for SonarQube Static Analysis Violations
- A Software-Repair Robot based on Continual Learning
- T5APR: Empowering Automated Program Repair across Languages through Checkpoint Ensemble
- A Comprehensive Study of Code-removal Patches in Automated Program Repair
- Test-based Patch Clustering for Automatically-Generated Patches Assessment
- Augmenting Diffs With Runtime Information
- STEAM: Simulating the InTeractive BEhavior of ProgrAMmers for Automatic Bug Fixing
- Risks of ignoring uncertainty propagation in AI-augmented security pipelines
- Estimating the Potential of Program Repair Search Spaces with Commit Analysis
- RePaCA: Leveraging Reasoning Large Language Models for Static Automated Patch Correctness Assessment
- MultiMend: Multilingual Program Repair with Context Augmentation and Multi-Hunk Patch Generation