1 paper · 1 filter
Jongyeop Hyun, Taeyoung Kim, Hyounghun Kim
Large Vision-Language Models produce fluent image descriptions but offer limited semantic control: users cannot reliably specify whether captions should emphasize attributes, relat…