30 citations · 47 across the 3 of their papers we have counts for
5 papers
COTS: Collaborative Two-Stream Vision-Language Pre-Training Model for Cross-Modal Retrieval
Haoyu Lu, Nanyi Fei, Yuqi Huo +3
Large-scale single-stream pre-training has shown dramatic performance in image-text retrieval. Regrettably, it faces low inference efficiency due to heavy attention layers. Recentl…
Pre-Trained Models: Past, Present and Future
Xu Han, Zhengyan Zhang, Ning Ding +21
Large-scale pre-trained models (PTMs) such as BERT and GPT have recently achieved great success and become a milestone in the field of artificial intelligence (AI). Owing to sophis…
WenLan: Bridging Vision and Language by Large-Scale Multi-Modal Pre-Training
Yuqi Huo, Manli Zhang, Guangzhen Liu +32
Multi-modal pre-training models have been intensively explored to bridge vision and language in recent years. However, most of them explicitly model the cross-modal interaction bet…
Learning Depth-Guided Convolutions for Monocular 3D Object Detection
Mingyu Ding, Yuqi Huo, Hongwei Yi +4
3D object detection from a single image without LiDAR is a challenging task due to the lack of accurate depth information. Conventional 2D convolutions are unsuitable for this task…
Mobile Video Action Recognition
Yuqi Huo, Xiaoli Xu, Yao Lu +3
Video action recognition, which is topical in computer vision and video analysis, aims to allocate a short video clip to a pre-defined category such as brushing hair or climbing st…