20 citations · 33 across the 8 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2023
TeachCLIP: Multi-Grained Teaching for Efficient Text-to-Video Retrieval
Kaibin Tian, Ruixiang Zhao, Hu Hu +4
For text-to-video retrieval (T2VR), which aims to retrieve unlabeled videos by ad-hoc textual queries, CLIP-based methods are dominating. Compared to CLIP4Clip which is efficient a…
cs.CV2020
STH: Spatio-Temporal Hybrid Convolution for Efficient Action Recognition
Xu Li, Jingwen Wang, Lin Ma +4
Effective and Efficient spatio-temporal modeling is essential for action recognition. Existing methods suffer from the trade-off between model performance and model complexity. In…