activity
20212023
most citedMultimodal Open-Vocabulary Video Classification via Pre-Trained Vision and Language Models

9 citations · 44 across the 10 of their papers we have counts for

collaborators

10 papers

cs.CV20237 cited

Taming Diffusion Models for Audio-Driven Co-Speech Gesture Generation

Lingting Zhu, Xian Liu, Xuanyu Liu +3

Animating virtual avatars to make co-speech gestures facilitates various applications in human-machine interaction. The existing methods mainly rely on generative adversarial netwo…

cs.CV20228 cited

Motion-inductive Self-supervised Object Discovery in Videos

Shuangrui Ding, Weidi Xie, Yabo Chen +4

In this paper, we consider the task of unsupervised object discovery in videos. Previous works have shown promising results via processing optical flows to segment objects. However…

cs.CV20221 cited

Static and Dynamic Concepts for Self-supervised Video Representation Learning

Rui Qian, Shuangrui Ding, Xian Liu +1

In this paper, we propose a novel learning scheme for self-supervised video representation learning. Motivated by how humans understand videos, we propose to first learn general vi…

cs.CV20225 cited

Exploring Fine-Grained Audiovisual Categorization with the SSW60 Dataset

Grant Van Horn, Rui Qian, Kimberly Wilber +3

We present a new benchmark dataset, Sapsucker Woods 60 (SSW60), for advancing research on audiovisual fine-grained categorization. While our community has made great strides in fin…

cs.CV20229 cited

Multimodal Open-Vocabulary Video Classification via Pre-Trained Vision and Language Models

Rui Qian, Yeqing Li, Zheng Xu +3

Utilizing vision and language models (VLMs) pre-trained on large-scale image-text pairs is becoming a promising paradigm for open-vocabulary visual recognition. In this work, we ex…

cs.CV20221 cited

Controllable Augmentations for Video Representation Learning

Rui Qian, Weiyao Lin, John See +1

This paper focuses on self-supervised video representation learning. Most existing approaches follow the contrastive learning pipeline to construct positive and negative pairs by s…