A VLM Answer Is Not an Anomaly Score: Rank Compression Across Image and Video Anomaly Detection
arXiv:2608.21244
Abstract
Anomaly detection aims to identify observations that deviate from normal patterns. Recent work has increasingly used pretrained vision--language models (VLMs) for training-free image and video anomaly detection without task-specific retraining. Anomaly detection is commonly evaluated by how well anomaly scores rank anomalous images or video frames above normal ones. Generative VLMs, however, assign probabilities to possible answers and then decode a single answer. This decoding step may discard ordering information. We call this loss decoded-answer rank compression. To isolate this effect, we compare two ways of scoring the same VLM output: one uses only the decoded answer, while the other computes a probability-weighted score over all possible answers. Probability-weighted scoring consistently outperforms decoded-answer scoring across image and video anomaly detection benchmarks, using different VLMs and answer scales. The mean gains range from 7.66 to 19.95 points on the primary metrics. Most of this gap comes from decoded-answer ties. Breaking decoded-answer ties with answer probabilities recovers at least 95% of the average performance gap on every benchmark. We further investigate how these ties affect evaluation. We find that evaluating tied scores one input at a time makes the reported AUROC depend on input order, and that linear PR interpolation can inflate reported performance. These results reveal that both how VLM outputs are scored and how discrete anomaly scores are evaluated affect reported anomaly-detection performance.
Preprint