1 paper
Jongyeop Hyun, Taeyoung Kim, Hyounghun Kim
Large Vision-Language Models produce fluent image descriptions but offer limited semantic control: users cannot reliably specify whether captions should emphasize attributes, relat…