Fooling OCR Systems with Adversarial Text Images
arXiv:1802.05385
Abstract
We demonstrate that state-of-the-art optical character recognition (OCR) based on deep learning is vulnerable to adversarial images. Minor modifications to images of printed text, which do not change the meaning of the text to a human reader, cause the OCR system to "recognize" a different text where certain words chosen by the adversary are replaced by their semantic opposites. This completely changes the meaning of the output produced by the OCR system and by the NLP applications that use OCR for preprocessing their inputs.
References in corpus (6)
- Obfuscated Gradients Give a False Sense of Security: Circumventing Defenses to Adversarial Examples
- Delving into Transferable Adversarial Examples and Black-box Attacks
- Generating Natural Adversarial Examples
- Countering Adversarial Images using Input Transformations
- Towards Crafting Text Adversarial Samples
- Deceiving Google's Perspective API Built for Detecting Toxic Comments
Cited by in corpus (7)
- Adversarial Attacks on Deep Learning Models in Natural Language Processing: A Survey
- Generating Natural Language Adversarial Examples on a Large Scale with Generative Models
- Universal Defensive Underpainting Patch: Making Your Text Invisible to Optical Character Recognition
- Black-box Adversarial Attacks on Network-wide Multi-step Traffic State Prediction Models
- Simple Transparent Adversarial Examples
- Adversarial Attacks on Binary Image Recognition Systems
- FAWA: Fast Adversarial Watermark Attack on Optical Character Recognition (OCR) Systems