27 citations · 33 across the 3 of their papers we have counts for
3 papers
cs.CV2023★ 4 cited
Fine-grained Text-Video Retrieval with Frozen Image Encoders
Zuozhuo Dai, Fangtao Shao, Qingkun Su +2
State-of-the-art text-video retrieval (TVR) methods typically utilize CLIP and cosine similarity for efficient retrieval. Meanwhile, cross attention methods, which employ a transfo…
cs.CV2023★ 2 cited
Towards Robust Video Instance Segmentation with Temporal-Aware Transformer
Zhenghao Zhang, Fangtao Shao, Zuozhuo Dai +1
Most existing transformer based video instance segmentation methods extract per frame features independently, hence it is challenging to solve the appearance deformation problem. I…
cs.CV2020★ 27 cited
Not only Look, but also Listen: Learning Multimodal Violence Detection under Weak Supervision
Peng Wu, Jing Liu, Yujia Shi +4
Violence detection has been studied in computer vision for years. However, previous work are either superficial, e.g., classification of short-clips, and the single scenario, or un…