4 papers · 1 filter
MMVIAD: Multi-view Multi-task Video Understanding for Industrial Anomaly Detection
Xiran Zhao, Jing Jin, Yan Bai +6
Industrial anomaly detection is critical for manufacturing quality control, yet existing datasets mainly focus on static images or sparse views, which do not fully reflect continuo…
Multi-turn Physics-informed Vision-language Model for Physics-grounded Anomaly Detection
Yao Gu, Xiaohao Xu, Yingna Wu
Vision-Language Models (VLMs) demonstrate strong general-purpose reasoning but remain limited in physics-grounded anomaly detection, where causal understanding of dynamics is essen…
Incremental Object Keypoint Learning
Mingfu Liang, Jiahuan Zhou, Xu Zou +1
Existing progress in object keypoint estimation primarily benefits from the conventional supervised learning paradigm based on numerous data labeled with pre-defined keypoints. How…
A Grammatical Compositional Model for Video Action Detection
Zhijun Zhang, Xu Zou, Jiahuan Zhou +2
Analysis of human actions in videos demands understanding complex human dynamics, as well as the interaction between actors and context. However, these interaction relationships us…