CoEdPilot: Recommending Code Edits with Learned Prior Edit Relevance, Project-wise Awareness, and Interactive Nature
arXiv:2408.01733 · doi:10.1145/3650212.3652142
Abstract
Recent years have seen the development of LLM-based code generation. Compared to generating code in a software project, incremental code edits are empirically observed to be more frequent. The emerging code editing approaches usually formulate the problem as generating an edit based on known relevant prior edits and context. However, practical code edits can be more complicated. First, an editing session can include multiple (ir)relevant edits to the code under edit. Second, the inference of the subsequent edits is non-trivial as the scope of its ripple effect can be the whole project. In this work, we propose CoEdPilot, an LLM-driven solution to recommend code edits by discriminating the relevant edits, exploring their interactive natures, and estimating its ripple effect in the project. Specifically, CoEdPilot orchestrates multiple neural transformers to identify what and how to edit in the project regarding both edit location and edit content. When a user accomplishes an edit with an optional editing description, a Subsequent Edit Analysis first reports the most relevant files in the project with what types of edits (e.g., keep, insert, and replace) can happen for each line of their code. Next, an Edit-content Generator generates concrete edit options for the lines of code, regarding its relevant prior changes reported by an Edit-dependency Analyzer. Lastly, both the Subsequent Edit Analysis and the Edit-content Generator capture relevant prior edits as feedback to readjust their recommendations. We train our models by collecting over 180K commits from 471 open-source projects in 5 programming languages. Our extensive experiments show that CoEdPilot can well predict the edits (i.e., predicting edit location with an accuracy of 70.8%-85.3%, and the edit content with an exact match rate of 41.8% and BLEU4 score of 60.7)...
13 pages, 7 figures
References in corpus (15)
- CURE: Code-Aware Neural Machine Translation for Automatic Program Repair
- StarCoder: may the source be with you!
- GraphCodeBERT: Pre-training Code Representations with Data Flow
- InCoder: A Generative Model for Code Infilling and Synthesis
- CODIT: Code Editing with Tree-Based Neural Models
- Prompting Large Language Model for Machine Translation: A Case Study
- CodeT: Code Generation with Generated Tests
- No More Manual Tests? Evaluating and Improving ChatGPT for Unit Test Generation
- CodeEditor: Learning to Edit Source Code with Pre-trained Models
- CodeT5+: Open Code Large Language Models for Code Understanding and Generation
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and Generation
- Reading Between the Lines: Modeling User Behavior and Costs in AI-Assisted Programming
- TreeGen: A Tree-Based Transformer Architecture for Code Generation
- Pre-Trained Models: Past, Present and Future
- CoditT5: Pretraining for Source Code and Natural Language Editing