most citedImageBERT: Cross-modal Pre-training with Large-scale Weak-supervised Image-Text Data

156 citations · 162 across the 4 of their papers we have counts for

collaborators

5 papers

cs.IR2021

Deep Keyphrase Completion

Yu Zhao, Jia Song, Huali Feng +4

Keyphrase provides accurate information of document content that is highly compact, concise, full of meanings, and widely used for discourse comprehension, organization, and text r…

cs.CV20214 cited

Learning Granularity-Aware Convolutional Neural Network for Fine-Grained Visual Classification

Jianwei Song, Ruoyu Yang

Locating discriminative parts plays a key role in fine-grained visual classification due to the high similarities between different objects. Recent works based on convolutional neu…

cs.CV2021

Feature Boosting, Suppression, and Diversification for Fine-Grained Visual Classification

Jianwei Song, Ruoyu Yang

Learning feature representation from discriminative local regions plays a key role in fine-grained visual classification. Employing attention mechanisms to extract part features ha…

cs.CV20202 cited

A Novel Video Salient Object Detection Method via Semi-supervised Motion Quality Perception

Chenglizhao Chen, Jia Song, Chong Peng +2

Previous video salient object detection (VSOD) approaches have mainly focused on designing fancy networks to achieve their performance improvements. However, with the slow-down in…

cs.CV2020156 cited

ImageBERT: Cross-modal Pre-training with Large-scale Weak-supervised Image-Text Data

Di Qi, Lin Su, Jia Song +3

In this paper, we introduce a new vision-language pre-trained model -- ImageBERT -- for image-text joint embedding. Our model is a Transformer-based model, which takes different mo…