6 papers · 1 filter
Learnable Expansion of Graph Operators for Multi-Modal Feature Fusion
Dexuan Ding, Lei Wang, Liyun Zhu +2
In computer vision tasks, features often come from diverse representations, domains (e.g., indoor and outdoor), and modalities (e.g., text, images, and videos). Effectively fusing…
When Spatial meets Temporal in Action Recognition
Huilin Chen, Lei Wang, Yifan Chen +2
Video action recognition has made significant strides, but challenges remain in effectively using both spatial and temporal information. While existing methods often focus on eithe…
Motion meets Attention: Video Motion Prompts
Qixiang Chen, Lei Wang, Piotr Koniusz +1
Videos contain rich spatio-temporal information. Traditional methods for extracting motion, used in tasks such as action recognition, often rely on visual contents rather than prec…
Adaptive Multi-head Contrastive Learning
Lei Wang, Piotr Koniusz, Tom Gedeon +1
In contrastive learning, two views of an original image, generated by different augmentations, are considered a positive pair, and their similarity is required to be high. Similarl…
TrackNetV4: Enhancing Fast Sports Object Tracking with Motion Attention Maps
Arjun Raj, Lei Wang, Tom Gedeon
Accurately detecting and tracking high-speed, small objects, such as balls in sports videos, is challenging due to factors like motion blur and occlusion. Although recent deep lear…
Taylor Videos for Action Recognition
Lei Wang, Xiuyuan Yuan, Tom Gedeon +1
Effectively extracting motions from video is a critical and long-standing problem for action recognition. This problem is very challenging because motions (i) do not have an explic…