VulBERTa: Simplified Source Code Pre-Training for Vulnerability Detection
arXiv:2205.12424 · doi:10.1109/IJCNN55064.2022.9892280
Abstract
This paper presents VulBERTa, a deep learning approach to detect security vulnerabilities in source code. Our approach pre-trains a RoBERTa model with a custom tokenisation pipeline on real-world code from open-source C/C++ projects. The model learns a deep knowledge representation of the code syntax and semantics, which we leverage to train vulnerability detection classifiers. We evaluate our approach on binary and multi-class vulnerability detection tasks across several datasets (Vuldeepecker, Draper, REVEAL and muVuldeepecker) and benchmarks (CodeXGLUE and D2A). The evaluation results show that VulBERTa achieves state-of-the-art performance and outperforms existing approaches across different datasets, despite its conceptual simplicity, and limited cost in terms of size of training data and number of model parameters.
Accepted as a conference paper at IJCNN 2022
References in corpus (7)
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- VulDeePecker: A Deep Learning-Based System for Vulnerability Detection
- CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation
- Exploring Software Naturalness through Neural Language Models
- Software Vulnerability Detection via Deep Learning over Disaggregated Code Graph Representation
- On using distributed representations of source code for the detection of C security vulnerabilities
- CoTexT: Multi-task Learning with Code-Text Transformer
Cited by in corpus (7)
- A Survey of Source Code Representations for Machine Learning-Based Cybersecurity Tasks
- An Unbiased Transformer Source Code Learning with Semantic Vulnerability Graph
- SecureBERT 2.0: Advanced Language Model for Cybersecurity Intelligence
- MalCodeAI: Autonomous Vulnerability Detection and Remediation via Language Agnostic Code Reasoning
- Nested Dirichlet models for unsupervised attack pattern detection in honeypot data
- Similar but Patched Code Considered Harmful -- The Impact of Similar but Patched Code on Recurring Vulnerability Detection and How to Remove Them
- LLaVul: A Multimodal LLM for Interpretable Vulnerability Reasoning about Source Code