2 citations · 2 across the 2 of their papers we have counts for
2 papers
cs.CV2024★ 2 cited
Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP
Yating Yu, Congqi Cao, Yueran Zhang +3
Zero-shot action recognition (ZSAR) requires collaborative multi-modal spatiotemporal understanding. However, finetuning CLIP directly for ZSAR yields suboptimal performance, given…
cs.CV2024
Task-Adapter: Task-specific Adaptation of Image Models for Few-shot Action Recognition
Congqi Cao, Yueran Zhang, Yating Yu +3
Existing works in few-shot action recognition mostly fine-tune a pre-trained image model and design sophisticated temporal alignment modules at feature level. However, simply fully…