HiProbe-VAD: Video Anomaly Detection via Hidden States Probing in Tuning-Free Multimodal LLMs
arXiv:2507.17394 · doi:10.1145/3746027.3755575
Abstract
Video Anomaly Detection (VAD) aims to identify and locate deviations from normal patterns in video sequences. Traditional methods often struggle with substantial computational demands and a reliance on extensive labeled datasets, thereby restricting their practical applicability. To address these constraints, we propose HiProbe-VAD, a novel framework that leverages pre-trained Multimodal Large Language Models (MLLMs) for VAD without requiring fine-tuning. In this paper, we discover that the intermediate hidden states of MLLMs contain information-rich representations, exhibiting higher sensitivity and linear separability for anomalies compared to the output layer. To capitalize on this, we propose a Dynamic Layer Saliency Probing (DLSP) mechanism that intelligently identifies and extracts the most informative hidden states from the optimal intermediate layer during the MLLMs reasoning. Then a lightweight anomaly scorer and temporal localization module efficiently detects anomalies using these extracted hidden states and finally generate explanations. Experiments on the UCF-Crime and XD-Violence datasets demonstrate that HiProbe-VAD outperforms existing training-free and most traditional approaches. Furthermore, our framework exhibits remarkable cross-model generalization capabilities in different MLLMs without any tuning, unlocking the potential of pre-trained MLLMs for video anomaly detection and paving the way for more practical and scalable solutions.
Accepted by ACM MM 2025
References in corpus (22)
- Sparks of Artificial General Intelligence: Early experiments with GPT-4
- Qwen2.5-VL Technical Report
- Anomaly Locality in Video Surveillance
- LLaVA-OneVision: Easy Visual Task Transfer
- How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
- Qwen2.5-Omni Technical Report
- Dual Memory Units with Uncertainty Regulation for Weakly Supervised Video Anomaly Detection
- Does Representation Matter? Exploring Intermediate Layers in Large Language Models
- How Large Language Models Encode Context Knowledge? A Layer-Wise Probing Study
- Video Anomaly Detection and Explanation via Large Language Models
- LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations
- Text Prompt with Normality Guidance for Weakly Supervised Video Anomaly Detection
- Unsupervised Video Anomaly Detection with Diffusion Models Conditioned on Compact Motion Representations
- GlanceVAD: Exploring Glance Supervision for Label-efficient Video Anomaly Detection
- Layer by Layer: Uncovering Hidden Representations in Language Models
- Exploring Information Processing in Large Language Models: Insights from Information Bottleneck Theory
- Entropy-Lens: Uncovering Decision Strategies in LLMs
- Is Bigger and Deeper Always Better? Probing LLaMA Across Scales and Layers
- Talking Heads: Understanding Inter-layer Communication in Transformer Language Models
- VERA: Explainable Video Anomaly Detection via Verbalized Learning of Vision-Language Models
- Holmes-VAU: Towards Long-term Video Anomaly Understanding at Any Granularity
- Investigating Layer Importance in Large Language Models