Practical Program Repair in the Era of Large Pre-trained Language Models
arXiv:2210.14179 · doi:10.1109/ICSE48619.2023.00129
Abstract
Automated Program Repair (APR) aims to help developers automatically patch software bugs. However, current state-of-the-art traditional and learning-based APR techniques face the problem of limited patch variety, failing to fix complicated bugs. This is mainly due to the reliance on bug-fixing datasets to craft fix templates or directly predict potential patches. Large Pre-Trained Language Models (PLMs), trained using billions of text/code tokens, can potentially help avoid this issue. Very recently, researchers have directly leveraged PLMs for APR without relying on any bug-fixing datasets. Meanwhile, such existing work either failed to include state-of-the-art PLMs or was not evaluated on realistic datasets. In this work, we perform the first extensive study on directly applying PLMs for APR. We select 9 recent state-of-the-art PLMs, including both generative and infilling models, ranging from 125M to 20B in size. We designed 3 different repair settings to evaluate the different ways we can use PLMs to generate patches. We apply the PLMs under these repair settings on 5 datasets across 3 different languages and compare different PLMs in the number of bugs fixed, generation speed and compilation rate. Our study demonstrates that directly applying state-of-the-art PLMs can already substantially outperform all existing APR techniques on all our datasets. Among the studied PLMs, the scaling effect exists for APR where larger models tend to achieve better performance. Also, we show for the first time that suffix code after the buggy line (adopted in infilling-style APR) is important in not only generating more fixes but more patches with higher compilation rate. Besides patch generation, the PLMs consider correct patches to be more natural than other ones, and can even be leveraged for effective patch ranking or patch correctness checking.
References in corpus (21)
- Sequence to Sequence Learning with Neural Networks
- Scaling Laws for Neural Language Models
- Evaluating Large Language Models Trained on Code
- The Curious Case of Neural Text Degeneration
- The Pile: An 800GB Dataset of Diverse Text for Language Modeling
- CodeSearchNet Challenge: Evaluating the State of Semantic Code Search
- CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation
- SequenceR: Sequence-to-Sequence Learning for End-to-End Program Repair
- TBar: Revisiting Template-based Automated Program Repair
- CURE: Code-Aware Neural Machine Translation for Automatic Program Repair
- Less Training, More Repairing Please: Revisiting Automated Program Repair via Zero-shot Learning
- FixMiner: Mining Relevant Fix Patterns for Automated Program Repair
- Neural Program Repair with Execution-based Backpropagation
- Automatic Repair of Buggy If Conditions and Missing Preconditions with SMT
- InCoder: A Generative Model for Code Infilling and Synthesis
- TOGA: A Neural Method for Test Oracle Generation
- On Learning Meaningful Assert Statements for Unit Test Cases
- CM3: A Causal Masked Multimodal Model of the Internet
- Conversational Automated Program Repair
- DeepDebug: Fixing Python Bugs Using Stack Traces, Backtranslation, and Code Skeletons
- Attention: Not Just Another Dataset for Patch-Correctness Checking
Cited by in corpus (33)
- Keep the Conversation Going: Fixing 162 out of 337 bugs for $0.42 each using ChatGPT
- Copiloting the Copilots: Fusing Large Language Models with Completion Engines for Automated Program Repair
- Breaking the Silence: the Threats of Using LLMs in Software Engineering
- What Makes Good In-context Demonstrations for Code Intelligence Tasks with LLMs?
- Fixing Hardware Security Bugs with Large Language Models
- Generative AI for Self-Adaptive Systems: State of the Art and Research Roadmap
- The Future of AI-Driven Software Engineering
- Look Before You Leap: An Exploratory Study of Uncertainty Measurement for Large Language Models
- ITER: Iterative Neural Repair for Multi-Location Patches
- DevGPT: Studying Developer-ChatGPT Conversations
- KernelGPT: Enhanced Kernel Fuzzing via Large Language Models
- Fuzzing Automatic Differentiation in Deep-Learning Libraries
- Demystifying RCE Vulnerabilities in LLM-Integrated Apps
- Revisiting the Plastic Surgery Hypothesis via Large Language Models
- Generative AI for Pull Request Descriptions: Adoption, Impact, and Developer Interventions
- RepairLLaMA: Efficient Representations and Fine-Tuned Adapters for Program Repair
- FlakyFix: Using Large Language Models for Predicting Flaky Test Fix Categories and Test Code Repair
- PyTy: Repairing Static Type Errors in Python
- Isolating Compiler Bugs by Generating Effective Witness Programs with Large Language Models
- LPR: Large Language Models-Aided Program Reduction
- Towards Integrating Emerging AI Applications in SE Education
- Large Language Model Powered Symbolic Execution
- LogUpdater: Automated Detection and Repair of Specific Defects in Logging Statements
- STEAM: Simulating the InTeractive BEhavior of ProgrAMmers for Automatic Bug Fixing
- Exploring the Effectiveness of Abstract Syntax Tree Patterns for Algorithm Recognition
- Morescient GAI for Software Engineering (Extended Version)
- Thinking Machines: Mathematical Reasoning in the Age of LLMs
- Hierarchical Knowledge Injection for Improving LLM-based Program Repair
- Toward Non-Expert Customized Congestion Control: Large Language Model-Assisted CCA Code Generation with eBPF Deployment
- Seeing is Fixing: Cross-Modal Reasoning with Multimodal LLMs for Visual Software Issue Fixing
- MultiMend: Multilingual Program Repair with Context Augmentation and Multi-Hunk Patch Generation
- Out of Context: How important is Local Context in Neural Program Repair?
- Assessing the Latent Automated Program Repair Capabilities of Large Language Models using Round-Trip Translation