Revisiting the Plastic Surgery Hypothesis via Large Language Models
arXiv:2303.10494 · doi:10.1109/ASE56229.2023.00047
Abstract
Automated Program Repair (APR) aspires to automatically generate patches for an input buggy program. Traditional APR tools typically focus on specific bug types and fixes through the use of templates, heuristics, and formal specifications. However, these techniques are limited in terms of the bug types and patch variety they can produce. As such, researchers have designed various learning-based APR tools with recent work focused on directly using Large Language Models (LLMs) for APR. While LLM-based APR tools are able to achieve state-of-the-art performance on many repair datasets, the LLMs used for direct repair are not fully aware of the project-specific information such as unique variable or method names. The plastic surgery hypothesis is a well-known insight for APR, which states that the code ingredients to fix the bug usually already exist within the same project. Traditional APR tools have largely leveraged the plastic surgery hypothesis by designing manual or heuristic-based approaches to exploit such existing code ingredients. However, as recent APR research starts focusing on LLM-based approaches, the plastic surgery hypothesis has been largely ignored. In this paper, we ask the following question: How useful is the plastic surgery hypothesis in the era of LLMs? Interestingly, LLM-based APR presents a unique opportunity to fully automate the plastic surgery hypothesis via fine-tuning and prompting. To this end, we propose FitRepair, which combines the direct usage of LLMs with two domain-specific fine-tuning strategies and one prompting strategy for more powerful APR. Our experiments on the widely studied Defects4j 1.2 and 2.0 datasets show that FitRepair fixes 89 and 44 bugs (substantially outperforming the best-performing baseline by 15 and 8), respectively, demonstrating a promising future of the plastic surgery hypothesis in the era of LLMs.
References in corpus (25)
- Adam: A Method for Stochastic Optimization
- Sequence to Sequence Learning with Neural Networks
- PaLM: Scaling Language Modeling with Pathways
- Evaluating Large Language Models Trained on Code
- CodeSearchNet Challenge: Evaluating the State of Semantic Code Search
- CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation
- SequenceR: Sequence-to-Sequence Learning for End-to-End Program Repair
- Fine-Tuning Language Models from Human Preferences
- Practical Program Repair in the Era of Large Pre-trained Language Models
- TBar: Revisiting Template-based Automated Program Repair
- CURE: Code-Aware Neural Machine Translation for Automatic Program Repair
- CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis
- Less Training, More Repairing Please: Revisiting Automated Program Repair via Zero-shot Learning
- Elixir: Effective object-oriented program repair
- FixMiner: Mining Relevant Fix Patterns for Automated Program Repair
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation
- Neural Program Repair with Execution-based Backpropagation
- Learning and Evaluating Contextual Embedding of Source Code
- Automatic Repair of Buggy If Conditions and Missing Preconditions with SMT
- InCoder: A Generative Model for Code Infilling and Synthesis
- SelfAPR: Self-supervised Program Repair with Test Execution Diagnostics
- Conversational Automated Program Repair
- ClassEval: A Manually-Crafted Benchmark for Evaluating LLMs on Class-level Code Generation
- Large Language Models are Edge-Case Fuzzers: Testing Deep Learning Libraries via FuzzGPT
- Evaluating Instruction-Tuned Large Language Models on Code Comprehension and Generation
Cited by in corpus (5)
- Generative AI for Self-Adaptive Systems: State of the Art and Research Roadmap
- RepairLLaMA: Efficient Representations and Fine-Tuned Adapters for Program Repair
- Hierarchical Knowledge Injection for Improving LLM-based Program Repair
- Seeing is Fixing: Cross-Modal Reasoning with Multimodal LLMs for Visual Software Issue Fixing
- MultiMend: Multilingual Program Repair with Context Augmentation and Multi-Hunk Patch Generation