210 citations · 316 across the 7 of their papers we have counts for
12 papers
MOFI: Learning Image Representations from Noisy Entity Annotated Images
Wentao Wu, Aleksei Timofeev, Chen Chen +8
We present MOFI, Manifold OF Images, a new vision foundation model designed to learn image representations from noisy entity annotated images. MOFI differs from previous work in tw…
Less is More: Removing Text-regions Improves CLIP Training Efficiency and Robustness
Liangliang Cao, Bowen Zhang, Chen Chen +5
The CLIP (Contrastive Language-Image Pre-training) model and its variants are becoming the de facto backbone in many applications. However, training a CLIP model from hundreds of m…
Auto-scaling Vision Transformers without Training
Wuyang Chen, Wei Huang, Xianzhi Du +3
This work targets automated designing and scaling of Vision Transformers (ViTs). The motivation comes from two pain spots: 1) the lack of efficient and principled methods for desig…
Revisiting 3D ResNets for Video Recognition
Xianzhi Du, Yeqing Li, Yin Cui +3
A recent work from Bello shows that training and scaling strategies may be more significant than model architectures for visual recognition. This short note studies effective train…
Simple Training Strategies and Model Scaling for Object Detection
Xianzhi Du, Barret Zoph, Wei-Chih Hung +1
The speed-accuracy Pareto curve of object detection systems have advanced through a combination of better model architectures, training and inference methods. In this paper, we met…
Dilated SpineNet for Semantic Segmentation
Abdullah Rashwan, Xianzhi Du, Xiaoqi Yin +1
Scale-permuted networks have shown promising results on object bounding box detection and instance segmentation. Scale permutation and cross-scale fusion of features enable the net…