Foundation Models and Transformers for Anomaly Detection: A Survey
arXiv:2507.15905 · doi:10.1016/j.inffus.2025.103517
Abstract
In line with the development of deep learning, this survey examines the transformative role of Transformers and foundation models in advancing visual anomaly detection (VAD). We explore how these architectures, with their global receptive fields and adaptability, address challenges such as long-range dependency modeling, contextual modeling and data scarcity. The survey categorizes VAD methods into reconstruction-based, feature-based and zero/few-shot approaches, highlighting the paradigm shift brought about by foundation models. By integrating attention mechanisms and leveraging large-scale pre-training, Transformers and foundation models enable more robust, interpretable, and scalable anomaly detection solutions. This work provides a comprehensive review of state-of-the-art techniques, their strengths, limitations, and emerging trends in leveraging these architectures for VAD.
References in corpus (63)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Distilling the Knowledge in a Neural Network
- Training language models to follow instructions with human feedback
- Learning to Prompt for Vision-Language Models
- On the Opportunities and Risks of Foundation Models
- A Survey of Large Language Models
- NICE: Non-linear Independent Components Estimation
- Flamingo: a Visual Language Model for Few-Shot Learning
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision
- Deep Learning for Anomaly Detection: A Survey
- Zero-Shot Text-to-Image Generation
- DINOv2: Learning Robust Visual Features without Supervision
- LAION-5B: An open large-scale dataset for training next generation image-text models
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
- Barlow Twins: Self-Supervised Learning via Redundancy Reduction
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
- Deep Industrial Image Anomaly Detection: A Survey
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- P-Tuning v2: Prompt Tuning Can Be Comparable to Fine-tuning Universally Across Scales and Tasks
- Omni-frequency Channel-selection Representations for Unsupervised Anomaly Detection
- Anomaly Detection in Video Using Predictive Convolutional Long Short-Term Memory Networks
- The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)
- Faster Segment Anything: Towards Lightweight SAM for Mobile Applications
- A Survey on Unsupervised Anomaly Detection Algorithms for Industrial Images
- Customized Segment Anything Model for Medical Image Segmentation
- Self-Supervised Masked Convolutional Transformer Block for Anomaly Detection
- A Complete Survey on Generative AI (AIGC): Is ChatGPT from GPT-4 to GPT-5 All You Need?
- Convolutional Transformer based Dual Discriminator Generative Adversarial Networks for Video Anomaly Detection
- Self-Supervised Anomaly Detection in Computer Vision and Beyond: A Survey and Outlook
- A Unified Model for Multi-class Anomaly Detection
- CogVLM: Visual Expert for Pretrained Language Models
- Large Language Models for Forecasting and Anomaly Detection: A Systematic Literature Review
- OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models
- Simple and Controllable Music Generation
- ConTNet: Why not use convolution and transformer at the same time?
- Unified Vision and Language Prompt Learning
- A Closer Look at the Explainability of Contrastive Language-Image Pre-training
- Masked Transformer for image Anomaly Localization
- AnomalyCLIP: Object-agnostic Prompt Learning for Zero-shot Anomaly Detection
- Mind the Pad -- CNNs can Develop Blind Spots
- Hierarchical Vector Quantized Transformer for Multi-class Unsupervised Anomaly Detection
- Reconstruction Student with Attention for Student-Teacher Pyramid Matching
- CHiLS: Zero-Shot Image Classification with Hierarchical Label Sets
- Towards Generic Anomaly Detection and Understanding: Large-scale Visual-linguistic Model (GPT-4V) Takes the Lead
- Multitask Vision-Language Prompt Tuning
- Exploring Plain ViT Reconstruction for Multi-class Unsupervised Anomaly Detection
- Generalizable Industrial Visual Anomaly Detection with Self-Induction Vision Transformer
- Bootstrap Fine-Grained Vision-Language Alignment for Unified Zero-Shot Anomaly Localization
- GPT-4V-AD: Exploring Grounding Potential of VQA-oriented GPT-4V for Zero-shot Anomaly Detection
- Siamese Transition Masked Autoencoders as Uniform Unsupervised Visual Anomaly Detector
- Random Word Data Augmentation with CLIP for Zero-Shot Anomaly Detection
- MuSc: Zero-Shot Industrial Anomaly Classification and Segmentation with Mutual Scoring of the Unlabeled Images
- CAINNFlow: Convolutional block Attention modules and Invertible Neural Networks Flow for anomaly detection and localization tasks
- Visual Anomaly Detection via Dual-Attention Transformer and Discriminative Flow
- Zero-Shot Anomaly Detection with Pre-trained Segmentation Models
- Generalized Out-of-Distribution Detection and Beyond in Vision Language Model Era: A Survey
- MMAD: A Comprehensive Benchmark for Multimodal Large Language Models in Industrial Anomaly Detection
- 2nd Place Winning Solution for the CVPR2023 Visual Anomaly and Novelty Detection Challenge: Multimodal Prompting for Data-centric Anomaly Detection
- A Survey on Diffusion Models for Anomaly Detection
- MAEDAY: MAE for few and zero shot AnomalY-Detection
- STNMamba: Mamba-based Spatial-Temporal Normality Learning for Video Anomaly Detection