24 citations · 56 across the 13 of their papers we have counts for
9 papers · 1 filter
Towards Generic Anomaly Detection and Understanding: Large-scale Visual-linguistic Model (GPT-4V) Takes the Lead
Yunkang Cao, Xiaohao Xu, Chen Sun +2
Anomaly detection is a crucial task across different domains and data types. However, existing anomaly detection models are often designed for specific domains and modalities. This…
Object-centric Video Representation for Long-term Action Anticipation
Ce Zhang, Changcheng Fu, Shijie Wang +4
This paper focuses on building object-centric representations for long-term action anticipation in videos. Our key motivation is that objects provide important cues to recognize an…
End-to-End Spatio-Temporal Action Localisation with Video Transformers
Alexey Gritsenko, Xuehan Xiong, Josip Djolonga +5
The most performant spatio-temporal action localisation models use external person proposals and complex external memory banks. We propose a fully end-to-end, purely-transformer ba…
Comparing Trajectory and Vision Modalities for Verb Representation
Dylan Ebert, Chen Sun, Ellie Pavlick
Three-dimensional trajectories, or the 3D position and rotation of objects over time, have been shown to encode key aspects of verb semantics (e.g., the meanings of roll vs. slide)…
Steerable Equivariant Representation Learning
Sangnie Bhardwaj, Willie McClinton, Tongzhou Wang +4
Pre-trained deep image representations are useful for post-training tasks such as classification through transfer learning, image retrieval, and object detection. Data augmentation…
A New Knowledge Distillation Network for Incremental Few-Shot Surface Defect Detection
Chen Sun, Liang Gao, Xinyu Li +1
Surface defect detection is one of the most essential processes for industrial quality inspection. Deep learning-based surface defect detection methods have shown great potential.…