Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
RegionReasoner: Region-Grounded Multi-Round Visual Reasoning
Wenfang Sun, Hao Chen, Yingjun Du +2
Large vision-language models have achieved remarkable progress in visual reasoning, yet most existing systems rely on single-step or text-only reasoning, limiting their ability to…
cs.CV2025
QUOTA: Quantifying Objects with Text-to-Image Models for Any Domain
Wenfang Sun, Yingjun Du, Gaowen Liu +2
We tackle the problem of quantifying the number of objects by a generative text-to-image model. Rather than retraining such a model for each new image domain of interest, which lea…