1 paper · 1 filter
Vésteinn Snæbjarnarson, Kevin Du, Niklas Stoehr +4
When a vision-language model (VLM) is prompted to identify an entity depicted in an image, it may answer 'I see a conifer,' rather than the specific label 'norway spruce'. This rai…