Sorting and Transforming Program Repair Ingredients via Deep Learning Code Similarities
arXiv:1707.04742 · doi:10.1109/SANER.2019.8668043
Abstract
In the field of automated program repair, the redundancy assumption claims large programs contain the seeds of their own repair. However, most redundancy-based program repair techniques do not reason about the repair ingredients---the code that is reused to craft a patch. We aim to reason about the repair ingredients by using code similarities to prioritize and transform statements in a codebase for patch generation. Our approach, DeepRepair, relies on deep learning to reason about code similarities. Code fragments at well-defined levels of granularity in a codebase can be sorted according to their similarity to suspicious elements (i.e., code elements that contain suspicious statements) and statements can be transformed by mapping out-of-scope identifiers to similar identifiers in scope. We examined these new search strategies for patch generation with respect to effectiveness from the viewpoint of a software maintainer. Our comparative experiments were executed on six open-source Java projects including 374 buggy program revisions and consisted of 19,949 trials spanning 2,616 days of computation time. DeepRepair's search strategy using code similarities generally found compilable ingredients faster than the baseline, jGenProg, but this improvement neither yielded test-adequate patches in fewer attempts (on average) nor found significantly more patches than the baseline. Although the patch counts were not statistically different, there were notable differences between the nature of DeepRepair patches and baseline patches. The results demonstrate that our learning-based approach finds patches that cannot be found by existing redundancy-based repair techniques.
camera-ready paper for SANER 2019
References in corpus (6)
- Nopol: Automatic Repair of Conditional Statement Bugs in Java Programs
- Automatic Repair of Real Bugs in Java: A Large-Scale Experiment on the Defects4J Dataset
- Sorting and Transforming Program Repair Ingredients via Deep Learning Code Similarities
- Automated Correction for Syntax Errors in Programming Assignments using Recurrent Neural Networks
- Learning API Usages from Bytecode: A Statistical Approach
- ASTOR: Evolutionary Automatic Software Repair for Java
Cited by in corpus (35)
- SequenceR: Sequence-to-Sequence Learning for End-to-End Program Repair
- TBar: Revisiting Template-based Automated Program Repair
- Neural Machine Translation Inspired Binary Code Similarity Comparison beyond Function Pairs
- Sorting and Transforming Program Repair Ingredients via Deep Learning Code Similarities
- Checking Smart Contracts with Structural Code Embedding
- On the Efficiency of Test Suite based Program Repair: A Systematic Assessment of 16 Automated Repair Systems for Java Programs
- Perfection Not Required? Human-AI Partnerships in Code Translation
- Ultra-Large Repair Search Space with Automatically Mined Templates: the Cardumen Mode of Astor
- Deep Learning-based Software Engineering: Progress, Challenges, and Opportunities
- iFixR: Bug Report driven Program Repair
- Automated Patch Assessment for Program Repair at Scale
- Generating Question Titles for Stack Overflow from Mined Code Snippets
- A Literature Study of Embeddings on Source Code
- On the Replicability and Reproducibility of Deep Learning in Software Engineering
- SCELMo: Source Code Embeddings from Language Models
- Astor: Exploring the Design Space of Generate-and-Validate Program Repair beyond GenProg
- Technical Q&A Site Answer Recommendation via Question Boosting
- Machine Learning in Compiler Optimisation
- The Remarkable Role of Similarity in Redundancy-based Program Repair
- Learning Lenient Parsing & Typing via Indirect Supervision
- Toward a Theory of Causation for Interpreting Neural Code Models
- Learning to Synthesize Programs as Interpretable and Generalizable Policies
- Practical Program Repair via Preference-based Ensemble Strategy
- Unit Test Case Generation with Transformers and Focal Context
- A Survey on Deep Learning for Software Engineering
- A Cross-Architecture Instruction Embedding Model for Natural Language Processing-Inspired Binary Code Analysis
- Empirical Review of Java Program Repair Tools: A Large-Scale Experiment on 2,141 Bugs and 23,551 Repair Attempts
- STRATA: Simple, Gradient-Free Attacks for Models of Code
- Which Syntactic Capabilities Are Statistically Learned by Masked Language Models for Code?
- Estimating the Potential of Program Repair Search Spaces with Commit Analysis
- Combining Code Embedding with Static Analysis for Function-Call Completion
- A Controlled Experiment of Different Code Representations for Learning-Based Bug Repair
- When Automated Program Repair Meets Regression Testing -- An Extensive Study on 2 Million Patches
- Obstacles in Fully Automatic Program Repair: A survey
- A Recurrent Neural Network Based Patch Recommender for Linux Kernel Bugs