Malware Makeover: Breaking ML-based Static Analysis by Modifying Executable Bytes
arXiv:1912.09064 · doi:10.1145/3433210.3453086
Abstract
Motivated by the transformative impact of deep neural networks (DNNs) in various domains, researchers and anti-virus vendors have proposed DNNs for malware detection from raw bytes that do not require manual feature engineering. In this work, we propose an attack that interweaves binary-diversification techniques and optimization frameworks to mislead such DNNs while preserving the functionality of binaries. Unlike prior attacks, ours manipulates instructions that are a functional part of the binary, which makes it particularly challenging to defend against. We evaluated our attack against three DNNs in white- and black-box settings, and found that it often achieved success rates near 100%. Moreover, we found that our attack can fool some commercial anti-viruses, in certain cases with a success rate of 85%. We explored several defenses, both new and old, and identified some that can foil over 80% of our evasion attempts. However, these defenses may still be susceptible to evasion by attacks, and so we advocate for augmenting malware-detection systems with methods that do not rely on machine learning.
Code for transformations at https://github.com/pwwl/enhanced-binary-diversification. Presentation at https://dl.acm.org/doi/10.1145/3433210.3453086. An author of a related work [32] contacted us regarding our characterization of their defense (Sec 2.2). They point out that our attack is not within the stated scope of their defense, but agree their defense would be ineffective against our attack
References in corpus (20)
- Explaining and Harnessing Adversarial Examples
- Distributed Representations of Sentences and Documents
- Theoretically Principled Trade-off between Robustness and Accuracy
- Evasion Attacks against Machine Learning at Test Time
- Provable defenses against adversarial examples via the convex outer adversarial polytope
- Certified Adversarial Robustness via Randomized Smoothing
- Countering Adversarial Images using Input Transformations
- Generating Adversarial Malware Examples for Black-Box Attacks Based on GAN
- Microsoft Malware Classification Challenge
- Prior Convictions: Black-Box Adversarial Attacks with Bandits and Priors
- Adversarial Transformation Networks: Learning to Generate Adversarial Examples
- On Detecting Adversarial Perturbations
- Imperceptible, Robust, and Targeted Adversarial Examples for Automatic Speech Recognition
- Houdini: Fooling Deep Structured Prediction Models
- EMBER: An Open Dataset for Training Static PE Malware Machine Learning Models
- On the Robustness of the CVPR 2018 White-Box Adversarial Example Defenses
- Feature Denoising for Improving Adversarial Robustness
- Generation & Evaluation of Adversarial Examples for Malware Obfuscation
- Non-Negative Networks Against Adversarial Attacks
- Adversarial Binaries for Authorship Identification
Cited by in corpus (8)
- Practical Attacks on Machine Learning: A Case Study on Adversarial Windows Malware
- A Robust Defense against Adversarial Attacks on Deep Learning-based Malware Detectors via (De)Randomized Smoothing
- Machine Learning for Windows Malware Detection and Classification: Methods, Challenges and Ongoing Research
- SLIFER: Investigating Performance and Robustness of Malware Detection Pipelines
- Towards a Practical Defense against Adversarial Attacks on Deep Learning-based Malware Detectors via Randomized Smoothing
- Black-box Attacks Against Neural Binary Function Detection
- Tarallo: Evading Behavioral Malware Detectors in the Problem Space
- A Comparison of State-of-the-Art Techniques for Generating Adversarial Malware Binaries