EVIL: Exploiting Software via Natural Language
arXiv:2109.00279 · doi:10.1109/ISSRE52982.2021.00042
Abstract
Writing exploits for security assessment is a challenging task. The writer needs to master programming and obfuscation techniques to develop a successful exploit. To make the task easier, we propose an approach (EVIL) to automatically generate exploits in assembly/Python language from descriptions in natural language. The approach leverages Neural Machine Translation (NMT) techniques and a dataset that we developed for this work. We present an extensive experimental study to evaluate the feasibility of EVIL, using both automatic and manual analysis, and both at generating individual statements and entire exploits. The generated code achieved high accuracy in terms of syntactic and semantic correctness.
Paper accepted at the 32nd International Symposium on Software Reliability Engineering (ISSRE 2021)
References in corpus (8)
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- Language Models are Few-Shot Learners
- NLTK: The Natural Language Toolkit
- A parallel corpus of Python functions and documentation strings for automated code documentation and code generation
- Does BLEU Score Work for Code Migration?
- The Impact of Preprocessing on Arabic-English Statistical and Neural Machine Translation
- Translation Quality Assessment: A Brief Survey on Manual and Automatic Methods
- Shellcode_IA32: A Dataset for Automatic Shellcode Generation