1 paper · 1 filter
Asma Farajidizaji, Akash Gupta, Vatsal Raina
Vision-language models are increasingly used to generate image captions in specific styles, such as humor or romantic. However, these transformer-based models often struggle with t…