9 citations · 13 across the 2 of their papers we have counts for
2 papers
cs.CV2022★ 9 cited
Multimodal Open-Vocabulary Video Classification via Pre-Trained Vision and Language Models
Rui Qian, Yeqing Li, Zheng Xu +3
Utilizing vision and language models (VLMs) pre-trained on large-scale image-text pairs is becoming a promising paradigm for open-vocabulary visual recognition. In this work, we ex…
cs.CV2021★ 4 cited
Exploring Temporal Granularity in Self-Supervised Video Representation Learning
Rui Qian, Yeqing Li, Liangzhe Yuan +7
This work presents a self-supervised learning framework named TeG to explore Temporal Granularity in learning video representations. In TeG, we sample a long clip from a video and…