2 citations · 2 across the 1 of their papers we have counts for
3 papers
cs.CV2024
QUOTA: Quantifying Objects with Text-to-Image Models for Any Domain
Wenfang Sun, Yingjun Du, Gaowen Liu +2
We tackle the problem of quantifying the number of objects by a generative text-to-image model. Rather than retraining such a model for each new image domain of interest, which lea…
cs.LG2024
IPO: Interpretable Prompt Optimization for Vision-Language Models
Yingjun Du, Wenfang Sun, Cees G. M. Snoek
Pre-trained vision-language models like CLIP have remarkably adapted to various downstream tasks. Nonetheless, their performance heavily depends on the specificity of the input tex…
cs.CV2024★ 2 cited
Training-Free Semantic Segmentation via LLM-Supervision
Wenfang Sun, Yingjun Du, Gaowen Liu +2
Recent advancements in open vocabulary models, like CLIP, have notably advanced zero-shot classification and segmentation by utilizing natural language for class-specific embedding…