A Closer Look at the Robustness of Vision-and-Language Pre-trained Models
arXiv:2012.08673
Abstract
Large-scale pre-trained multimodal transformers, such as ViLBERT and UNITER, have propelled the state of the art in vision-and-language (V+L) research to a new level. Although achieving impressive performance on standard tasks, to date, it still remains unclear how robust these pre-trained models are. To investigate, we conduct a host of thorough evaluations on existing pre-trained models over 4 different types of V+L specific model robustness: (i) Linguistic Variation; (ii) Logical Reasoning; (iii) Visual Content Manipulation; and (iv) Answer Distribution Shift. Interestingly, by standard model finetuning, pre-trained V+L models already exhibit better robustness than many task-specific state-of-the-art methods. To further enhance model robustness, we propose Mango, a generic and efficient approach that learns a Multimodal Adversarial Noise GeneratOr in the embedding space to fool pre-trained V+L models. Differing from previous studies focused on one specific type of robustness, Mango is task-agnostic, and enables universal performance lift for pre-trained models over diverse tasks designed to evaluate broad aspects of robustness. Comprehensive experiments demonstrate that Mango achieves new state of the art on 7 out of 9 robustness benchmarks, surpassing existing methods by a significant margin. As the first comprehensive study on V+L robustness, this work puts robustness of pre-trained models into sharper focus, pointing new directions for future study.
References in corpus (9)
- Microsoft COCO Captions: Data Collection and Evaluation Server
- Theoretically Principled Trade-off between Robustness and Accuracy
- Visual Entailment: A Novel Task for Fine-Grained Image Understanding
- Adversarial Training for Large Neural Language Models
- Contrastive Visual-Linguistic Pretraining
- VIVO: Visual Vocabulary Pre-Training for Novel Object Captioning
- Behind the Scene: Revealing the Secrets of Pre-trained Vision-and-Language Models
- Semantic Equivalent Adversarial Data Augmentation for Visual Question Answering
- Unsupervised Vision-and-Language Pre-training Without Parallel Images and Captions
Cited by in corpus (6)
- Multimodal Co-learning: Challenges, Applications with Datasets, Recent Advances and Future Directions
- Scaling Up Vision-Language Pre-training for Image Captioning
- HateProof: Are Hateful Meme Detection Systems really Robust?
- VL-LTR: Learning Class-wise Visual-Linguistic Representation for Long-Tailed Visual Recognition
- X-GGM: Graph Generative Modeling for Out-of-Distribution Generalization in Visual Question Answering
- Discovering the Unknown Knowns: Turning Implicit Knowledge in the Dataset into Explicit Training Examples for Visual Question Answering