Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
Your other Left! Vision-Language Models Fail to Identify Relative Positions in Medical Images
Daniel Wolf, Heiko Hillenhagen, Billurvan Taskin +4
Clinical decision-making relies heavily on understanding relative positions of anatomical structures and anomalies. Therefore, for Vision-Language Models (VLMs) to be applicable in…
cs.CV2025
A Survey on Quality Metrics for Text-to-Image Generation
Sebastian Hartwig, Dominik Engel, Leon Sick +6
AI-based text-to-image models do not only excel at generating realistic images, they also give designers more and more fine-grained control over the image content. Consequently, th…