45 citations · 97 across the 11 of their papers we have counts for
8 papers · 1 filter
Motion meets Attention: Video Motion Prompts
Qixiang Chen, Lei Wang, Piotr Koniusz +1
Videos contain rich spatio-temporal information. Traditional methods for extracting motion, used in tasks such as action recognition, often rely on visual contents rather than prec…
Authentic Emotion Mapping: Benchmarking Facial Expressions in Real News
Qixuan Zhang, Zhifeng Wang, Yang Liu +4
In this paper, we present a novel benchmark for Emotion Recognition using facial landmarks extracted from realistic news videos. Traditional methods relying on RGB images are resou…
Meet JEANIE: a Similarity Measure for 3D Skeleton Sequences via Temporal-Viewpoint Alignment
Lei Wang, Jun Liu, Liang Zheng +2
Video sequences exhibit significant nuisance variations (undesired effects) of speed of actions, temporal locations, and subjects' poses, leading to temporal-viewpoint misalignment…
A Closer Look at the Robustness of Contrastive Language-Image Pre-Training (CLIP)
Weijie Tu, Weijian Deng, Tom Gedeon
Contrastive Language-Image Pre-training (CLIP) models have demonstrated remarkable generalization capabilities across multiple challenging distribution shifts. However, there is st…
Training with Product Digital Twins for AutoRetail Checkout
Yue Yao, Xinyu Tian, Zheng Tang +4
Automating the checkout process is important in smart retail, where users effortlessly pass products by hand through a camera, triggering automatic product detection, tracking, and…
Large-scale Training Data Search for Object Re-identification
Yue Yao, Huan Lei, Tom Gedeon +1
We consider a scenario where we have access to the target domain, but cannot afford on-the-fly training data annotation, and instead would like to construct an alternative training…