17 citations · 71 across the 42 of their papers we have counts for
11 papers · 2 filters
Quo Vadis, Anomaly Detection? LLMs and VLMs in the Spotlight
Xi Ding, Lei Wang
Video anomaly detection (VAD) has witnessed significant advancements through the integration of large language models (LLMs) and vision-language models (VLMs), addressing critical…
Do Language Models Understand Time?
Xi Ding, Lei Wang
Large language models (LLMs) have revolutionized video-based computer vision applications, including action recognition, anomaly detection, and video summarization. Videos inherent…
When Spatial meets Temporal in Action Recognition
Huilin Chen, Lei Wang, Yifan Chen +2
Video action recognition has made significant strides, but challenges remain in effectively using both spatial and temporal information. While existing methods often focus on eithe…
Learnable Expansion of Graph Operators for Multi-Modal Feature Fusion
Dexuan Ding, Lei Wang, Liyun Zhu +2
In computer vision tasks, features often come from diverse representations, domains (e.g., indoor and outdoor), and modalities (e.g., text, images, and videos). Effectively fusing…
TrackNetV4: Enhancing Fast Sports Object Tracking with Motion Attention Maps
Arjun Raj, Lei Wang, Tom Gedeon
Accurately detecting and tracking high-speed, small objects, such as balls in sports videos, is challenging due to factors like motion blur and occlusion. Although recent deep lear…
Motion meets Attention: Video Motion Prompts
Qixiang Chen, Lei Wang, Piotr Koniusz +1
Videos contain rich spatio-temporal information. Traditional methods for extracting motion, used in tasks such as action recognition, often rely on visual contents rather than prec…