2 papers
cs.CV2024
BodyMetric: Evaluating the Realism of Human Bodies in Text-to-Image Generation
Nefeli Andreou, Varsha Vivek, Ying Wang +5
Accurately generating images of human bodies from text remains a challenging problem for state of the art text-to-image models. Commonly observed body-related artifacts include ext…
cs.CV2024
Benchmarking Zero-Shot Recognition with Vision-Language Models: Challenges on Granularity and Specificity
Zhenlin Xu, Yi Zhu, Tiffany Deng +6
This paper presents novel benchmarks for evaluating vision-language models (VLMs) in zero-shot recognition, focusing on granularity and specificity. Although VLMs excel in tasks li…