SNIFF: Reverse Engineering of Neural Networks with Fault Attacks
arXiv:2002.11021 · doi:10.1109/TR.2021.3105697
Abstract
Neural networks have been shown to be vulnerable against fault injection attacks. These attacks change the physical behavior of the device during the computation, resulting in a change of value that is currently being computed. They can be realized by various fault injection techniques, ranging from clock/voltage glitching to application of lasers to rowhammer. In this paper we explore the possibility to reverse engineer neural networks with the usage of fault attacks. SNIFF stands for sign bit flip fault, which enables the reverse engineering by changing the sign of intermediate values. We develop the first exact extraction method on deep-layer feature extractor networks that provably allows the recovery of the model parameters. Our experiments with Keras library show that the precision error for the parameter recovery for the tested networks is less than with the usage of 64-bit floats, which improves the current state of the art by 6 orders of magnitude. Additionally, we discuss the protection techniques against fault injection attacks that can be applied to enhance the fault resistance.
Published in IEEE Transactions on Reliability
References in corpus (7)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Stealing Machine Learning Models via Prediction APIs
- Qualitatively characterizing neural network optimization problems
- Terminal Brain Damage: Exposing the Graceless Degradation in Deep Neural Networks Under Hardware Fault Attacks
- A Simple Explanation for the Existence of Adversarial Examples with Small Hamming Distance
- Targeted Attack against Deep Neural Networks via Flipping Limited Weight Bits
- Multiple Fault Attack on PRESENT with a Hardware Trojan Implementation in FPGA