activity
20182026
most citedImproving CLIP Training with Language Rewrites

35 citations · 100 across the 15 of their papers we have counts for

collaborators
Showing cs.CVShow all

19 papers · 1 filter

cs.CV2026

Visual General Intelligence: A White Paper

Hirokatsu Kataoka, Yoshihiro Fukuhara, Yonglong Tian +18

This paper reconsiders intelligence from a vision-centered perspective and examines whether intelligence emerging from visual experience and learning may provide a pathway toward A…

cs.CV2026

Beyond Correctness: Benchmarking and Aligning Response Behaviors in Hybrid-Thinking MLLMs

Xinming Wang, Weinong Wang, Hongming Yang +13

Hybrid-thinking multimodal large language models (MLLMs) allow a single model to alternate between deliberative thinking and latency-efficient non-thinking inference. Although thes…

cs.CV2025★ 2 cited

Vision-Language Models Do Not Understand Negation

Kumail Alhamoud, Shaden Alshammari, Yonglong Tian +4

Many practical vision-language applications require models that understand negation, e.g., when using natural language to retrieve images which contain certain objects but not othe…

cs.CV2023

Learning Vision from Models Rivals Learning Vision from Data

Yonglong Tian, Lijie Fan, Kaifeng Chen +3

We introduce SynCLR, a novel approach for learning visual representations exclusively from synthetic images and synthetic captions, without any real data. We synthesize a large dat…

cs.CV2023★ 1 cited

Scaling Laws of Synthetic Images for Model Training ... for Now

Lijie Fan, Kaifeng Chen, Dilip Krishnan +3

Recent significant advances in text-to-image models unlock the possibility of training vision systems using synthetic images, potentially overcoming the difficulty of collecting cu…

cs.CV2023★ 1 cited

Leveraging Unpaired Data for Vision-Language Generative Models via Cycle Consistency

Tianhong Li, Sangnie Bhardwaj, Yonglong Tian +6

Current vision-language generative models rely on expansive corpora of paired image-text data to attain optimal performance and generalization capabilities. However, automatically…