Neural Transfer Learning for Repairing Security Vulnerabilities in C Code
arXiv:2104.08308 · doi:10.1109/TSE.2022.3147265
Abstract
In this paper, we address the problem of automatic repair of software vulnerabilities with deep learning. The major problem with data-driven vulnerability repair is that the few existing datasets of known confirmed vulnerabilities consist of only a few thousand examples. However, training a deep learning model often requires hundreds of thousands of examples. In this work, we leverage the intuition that the bug fixing task and the vulnerability fixing task are related and that the knowledge learned from bug fixes can be transferred to fixing vulnerabilities. In the machine learning community, this technique is called transfer learning. In this paper, we propose an approach for repairing security vulnerabilities named VRepair which is based on transfer learning. VRepair is first trained on a large bug fix corpus and is then tuned on a vulnerability fix dataset, which is an order of magnitude smaller. In our experiments, we show that a model trained only on a bug fix corpus can already fix some vulnerabilities. Then, we demonstrate that transfer learning improves the ability to repair vulnerable C functions. We also show that the transfer learning model performs better than a model trained with a denoising task and fine-tuned on the vulnerability fixing task. To sum up, this paper shows that transfer learning works well for repairing security vulnerabilities in C compared to learning on a small dataset.
References in corpus (8)
- Sequence to Sequence Learning with Neural Networks
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- VulDeePecker: A Deep Learning-Based System for Vulnerability Detection
- CVEfixes: Automated Collection of Vulnerabilities and Their Fixes from Open-Source Software
- Graph-based, Self-Supervised Program Repair from Diagnostic Feedback
- A Software-Repair Robot based on Continual Learning
- A ground-truth dataset of real security patches
- SeqTrans: Automatic Vulnerability Fix via Sequence to Sequence Learning
Cited by in corpus (11)
- Large Language Models for Code: Security Hardening and Adversarial Testing
- How Effective Are Neural Networks for Fixing Security Vulnerabilities
- Deep Learning-based Software Engineering: Progress, Challenges, and Opportunities
- Fixing Hardware Security Bugs with Large Language Models
- Enhanced Automated Code Vulnerability Repair using Large Language Models
- RepairLLaMA: Efficient Representations and Fine-Tuned Adapters for Program Repair
- A Survey of Source Code Representations for Machine Learning-Based Cybersecurity Tasks
- PyTy: Repairing Static Type Errors in Python
- Self-Supervised Learning to Prove Equivalence Between Straight-Line Programs via Rewrite Rules
- A Closer Look at the Security Risks in the Rust Ecosystem
- Self-Bootstrapping Automated Program Repair: Using LLMs to Generate and Evaluate Synthetic Training Data for Bug Repair