4 citations · 6 across the 2 of their papers we have counts for
2 papers
cs.CV2023★ 4 cited
Fine-grained Text-Video Retrieval with Frozen Image Encoders
Zuozhuo Dai, Fangtao Shao, Qingkun Su +2
State-of-the-art text-video retrieval (TVR) methods typically utilize CLIP and cosine similarity for efficient retrieval. Meanwhile, cross attention methods, which employ a transfo…
cs.CV2023★ 2 cited
Towards Robust Video Instance Segmentation with Temporal-Aware Transformer
Zhenghao Zhang, Fangtao Shao, Zuozhuo Dai +1
Most existing transformer based video instance segmentation methods extract per frame features independently, hence it is challenging to solve the appearance deformation problem. I…