45 citations · 54 across the 3 of their papers we have counts for
8 papers
Cross-Modal Attention Consistency for Video-Audio Unsupervised Learning
Shaobo Min, Qi Dai, Hongtao Xie +3
Cross-modal correlation provides an inherent supervision for video unsupervised representation learning. Existing methods focus on distinguishing different video clips by visual an…
Self-Supervised Learning with Swin Transformers
Zhenda Xie, Yutong Lin, Zhuliang Yao +4
We are witnessing a modeling shift from CNN to Transformers in computer vision. In this work, we present a self-supervised learning approach called MoBY, with Vision Transformers a…
Temporal Action Detection with Multi-level Supervision
Baifeng Shi, Qi Dai, Judy Hoffman +3
Training temporal action detection in videos requires large amounts of labeled data, yet such annotation is expensive to collect. Incorporating unlabeled or weakly-labeled data to…
Informative Dropout for Robust Representation Learning: A Shape-bias Perspective
Baifeng Shi, Dinghuai Zhang, Qi Dai +3
Convolutional Neural Networks (CNNs) are known to rely more on local texture rather than global shape when making decisions. Recent work also indicates a close relationship between…
Weakly-Supervised Action Localization by Generative Attention Modeling
Baifeng Shi, Qi Dai, Yadong Mu +1
Weakly-supervised temporal action localization is a problem of learning an action localization model with only video-level action labeling available. The general framework largely…
Improving the Learning of Multi-column Convolutional Neural Network for Crowd Counting
Zhi-Qi Cheng, Jun-Xiu Li, Qi Dai +3
Tremendous variation in the scale of people/head size is a critical problem for crowd counting. To improve the scale invariance of feature representation, recent works extensively…