2 citations · 2 across the 4 of their papers we have counts for
4 papers
Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP
Yating Yu, Congqi Cao, Yueran Zhang +3
Zero-shot action recognition (ZSAR) requires collaborative multi-modal spatiotemporal understanding. However, finetuning CLIP directly for ZSAR yields suboptimal performance, given…
Task-Adapter: Task-specific Adaptation of Image Models for Few-shot Action Recognition
Congqi Cao, Yueran Zhang, Yating Yu +3
Existing works in few-shot action recognition mostly fine-tune a pre-trained image model and design sophisticated temporal alignment modules at feature level. However, simply fully…
A New Comprehensive Benchmark for Semi-supervised Video Anomaly Detection and Anticipation
Congqi Cao, Yue Lu, Peng Wang +1
Semi-supervised video anomaly detection (VAD) is a critical task in the intelligent surveillance system. However, an essential type of anomaly in VAD named scene-dependent anomaly…
Co-Occurrence Matters: Learning Action Relation for Temporal Action Localization
Congqi Cao, Yizhe Wang, Yue Lu +2
Temporal action localization (TAL) is a prevailing task due to its great application potential. Existing works in this field mainly suffer from two weaknesses: (1) They often negle…