5 citations · 9 across the 6 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2024
Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation
Jianyu Zhang, Li Zhang, Shijian Li
The visual understanding are often approached from 3 granular levels: image, patch and pixel. Visual Tokenization, trained by self-supervised reconstructive learning, compresses vi…
cs.CV2024
A Framework For Image Synthesis Using Supervised Contrastive Learning
Yibin Liu, Jianyu Zhang, Li Zhang +2
Text-to-image (T2I) generation aims at producing realistic images corresponding to text descriptions. Generative Adversarial Network (GAN) has proven to be successful in this task.…