17 citations · 18 across the 2 of their papers we have counts for
4 papers
Vision-Language Pre-Training for Boosting Scene Text Detectors
Sibo Song, Jianqiang Wan, Zhibo Yang +4
Recently, vision-language joint representation learning has proven to be highly effective in various scenarios. In this paper, we specifically adapt vision-language joint learning…
Deep Adaptive Temporal Pooling for Activity Recognition
Sibo Song, Ngai-Man Cheung, Vijay Chandrasekhar +1
Deep neural networks have recently achieved competitive accuracy for human activity recognition. However, there is room for improvement, especially in modeling long-term temporal i…
Defense Against Adversarial Attacks with Saak Transform
Sibo Song, Yueru Chen, Ngai-Man Cheung +1
Deep neural networks (DNNs) are known to be vulnerable to adversarial perturbations, which imposes a serious threat to DNN-based decision systems. In this paper, we propose to appl…
Truly Multi-modal YouTube-8M Video Classification with Video, Audio, and Text
Zhe Wang, Kingsley Kuan, Mathieu Ravaut +13
The YouTube-8M video classification challenge requires teams to classify 0.7 million videos into one or more of 4,716 classes. In this Kaggle competition, we placed in the top 3% o…