1 paper
Yiming Zhao, Guorong Li, Laiyun Qing +5
Open-world object counting leverages the robust text-image alignment of pre-trained vision-language models (VLMs) to enable counting of arbitrary categories in images specified by…