6 papers
TD-VAD: Breaking Visual Dependence in Video Anomaly Detection with Text-Driven Learning
Shuangqing Zhang, Lei-Lei Ma, Zhao Wang +5
Visual data is typically a prerequisite for training existing video anomaly detection (VAD) methods. However, obtaining sufficient annotated anomaly data for training is challengin…
OpenGround: Planning-based Online Perception for Open-World 3D Visual Grounding
Wenyuan Huang, Zhenyu Zhang, Zhao Wang +4
3D visual grounding aims to locate objects based on natural language descriptions in 3D scenes. Existing supervised methods are limited by generalization and recent zero-shot metho…
Global Logic and Local Search: Dual-Stream Multimodal In-Context Learning for Verifiable Industrial Anomaly Detection
Runzhi Deng, Yundi Hu, Yiming Zhong +5
Large Multimodal Models (LMMs) show strong few-shot generalization, but industrial anomaly detection remains difficult because defects are small, input resolution is limited, and t…
CL-Anomaly: Layer-Adaptive Mixture-of-Experts with Multimodal Large Language Model for Continual Learning in Anomaly Detection
Wen Dong, Zhao Wang, Shuangqing Zhang +5
Multimodal Large Language Models (MLLMs) excel in diverse vision tasks, but full-parameter retraining is computationally expensive as real-world knowledge evolves. Existing continu…
One-to-More: High-Fidelity Training-Free Anomaly Generation with Attention Control
Haoxiang Rao, Zhao Wang, Chenyang Si +4
Industrial anomaly detection (AD) is characterized by an abundance of normal images but a scarcity of anomalous ones. Although numerous few-shot anomaly synthesis methods have been…
ABounD: Adversarial Boundary-Driven Few-Shot Learning for Multi-Class Anomaly Detection
Runzhi Deng, Yundi Hu, Xinshuang Zhang +5
Few-shot multi-class industrial anomaly detection identifies diverse defects across multiple categories using a single unified model and limited normal samples. Although vision-lan…