2 papers
cs.CV2025
QUOTA: Quantifying Objects with Text-to-Image Models for Any Domain
Wenfang Sun, Yingjun Du, Gaowen Liu +2
We tackle the problem of quantifying the number of objects by a generative text-to-image model. Rather than retraining such a model for each new image domain of interest, which lea…
cs.LG2024
IPO: Interpretable Prompt Optimization for Vision-Language Models
Yingjun Du, Wenfang Sun, Cees G. M. Snoek
Pre-trained vision-language models like CLIP have remarkably adapted to various downstream tasks. Nonetheless, their performance heavily depends on the specificity of the input tex…