activity
20142024
most citedDeepID-Net: multi-stage and deformable deep convolutional neural networks for object detection

134 citations · 232 across the 7 of their papers we have counts for

collaborators

7 papers

cs.CV2024

Personalized Representation from Personalized Generation

Shobhita Sundaram, Julia Chae, Yonglong Tian +2

Modern vision models excel at general purpose downstream tasks. It is unclear, however, how they may be used for personalized vision tasks, which are both fine-grained and data-sca…

cs.CV2023

Learning Vision from Models Rivals Learning Vision from Data

Yonglong Tian, Lijie Fan, Kaifeng Chen +3

We introduce SynCLR, a novel approach for learning visual representations exclusively from synthetic images and synthetic captions, without any real data. We synthesize a large dat…

cs.CV20231 cited

Leveraging Unpaired Data for Vision-Language Generative Models via Cycle Consistency

Tianhong Li, Sangnie Bhardwaj, Yonglong Tian +6

Current vision-language generative models rely on expansive corpora of paired image-text data to attain optimal performance and generalization capabilities. However, automatically…

cs.LG202311 cited

PFGM++: Unlocking the Potential of Physics-Inspired Generative Models

Yilun Xu, Ziming Liu, Yonglong Tian +3

We introduce a new family of physics-inspired generative models termed PFGM++ that unifies diffusion models and Poisson Flow Generative Models (PFGM). These models realize generati…

cs.CV201479 cited

DeepID-Net: Deformable Deep Convolutional Neural Networks for Object Detection

Wanli Ouyang, Xiaogang Wang, Xingyu Zeng +8

In this paper, we propose deformable deep convolutional neural networks for generic object detection. This new deep learning object detection framework has innovations in multiple…

cs.CV20147 cited

Pedestrian Detection aided by Deep Learning Semantic Tasks

Yonglong Tian, Ping Luo, Xiaogang Wang +1

Deep learning methods have achieved great success in pedestrian detection, owing to its ability to learn features from raw pixels. However, they mainly capture middle-level represe…