Enhanced Automated Code Vulnerability Repair using Large Language Models
arXiv:2401.03741 · doi:10.1016/j.engappai.2024.109291
Abstract
This research addresses the complex challenge of automated repair of code vulnerabilities, vital for enhancing digital security in an increasingly technology-driven world. The study introduces a novel and efficient format for the representation of code modification, using advanced Large Language Models (LLMs) such as Code Llama and Mistral. These models, fine-tuned on datasets featuring C code vulnerabilities, significantly improve the accuracy and adaptability of automated code repair techniques. A key finding is the enhanced repair accuracy of these models when compared to previous methods such as VulRepair, which underscores their practical utility and efficiency. The research also offers a critical assessment of current evaluation metrics, such as perfect predictions, and their limitations in reflecting the true capabilities of automated repair models in real-world scenarios. Following this, it underscores the importance of using test datasets devoid of train samples, emphasizing the need for dataset integrity to enhance the effectiveness of LLMs in code repair tasks. The significance of this work is its contribution to digital security, setting new standards for automated code vulnerability repair and paving the way for future advancements in the fields of cybersecurity and artificial intelligence. The study does not only highlight the potential of LLMs in enhancing code security but also fosters further exploration and research in these crucial areas.
References in corpus (17)
- Training language models to follow instructions with human feedback
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- LoRA: Low-Rank Adaptation of Large Language Models
- PaLM: Scaling Language Modeling with Pathways
- Scaling Laws for Neural Language Models
- LaMDA: Language Models for Dialog Applications
- Nopol: Automatic Repair of Conditional Statement Bugs in Java Programs
- Using DeepSpeed and Megatron to Train Megatron-Turing NLG 530B, A Large-Scale Generative Language Model
- FixMiner: Mining Relevant Fix Patterns for Automated Program Repair
- GLaM: Efficient Scaling of Language Models with Mixture-of-Experts
- Neural Transfer Learning for Repairing Security Vulnerabilities in C Code
- InCoder: A Generative Model for Code Infilling and Synthesis
- GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
- UL2: Unifying Language Learning Paradigms
- AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
- CodeGen2: Lessons for Training LLMs on Programming and Natural Languages
- The case for 4-bit precision: k-bit Inference Scaling Laws