33 citations · 65 across the 7 of their papers we have counts for
7 papers
TinyCLIP: CLIP Distillation via Affinity Mimicking and Weight Inheritance
Kan Wu, Houwen Peng, Zhenghong Zhou +10
In this paper, we propose a novel cross-modal distillation method, called TinyCLIP, for large-scale language-image pre-trained models. The method introduces two core techniques: af…
Improving Visual Quality of Image Synthesis by A Token-based Generator with Transformers
Yanhong Zeng, Huan Yang, Hongyang Chao +2
We present a new perspective of achieving image synthesis by viewing this task as a visual token generation problem. Different from existing paradigms that directly synthesize a fu…
CoCo-BERT: Improving Video-Language Pre-training with Contrastive Cross-modal Matching and Denoising
Jianjie Luo, Yehao Li, Yingwei Pan +3
BERT-type structure has led to the revolution of vision-language pre-training and the achievement of state-of-the-art results on numerous vision-language downstream tasks. Existing…
CORE-Text: Improving Scene Text Detection with Contrastive Relational Reasoning
Jingyang Lin, Yingwei Pan, Rongfeng Lai +3
Localizing text instances in natural scenes is regarded as a fundamental challenge in computer vision. Nevertheless, owing to the extremely varied aspect ratios and scales of text…
Automatically Building Face Datasets of New Domains from Weakly Labeled Data with Pretrained Models
Shengyong Ding, Junyu Wu, Wei Xu +1
Training data are critical in face recognition systems. However, labeling a large scale face data for a particular domain is very tedious. In this paper, we propose a method to aut…
Deep Joint Face Hallucination and Recognition
Junyu Wu, Shengyong Ding, Wei Xu +1
Deep models have achieved impressive performance for face hallucination tasks. However, we observe that directly feeding the hallucinated facial images into recog- nition models ca…