1 paper
Wes Robbins, Zanyar Zohourianshahzadi, Jugal Kalita
Vision-language models can assess visual context in an image and generate descriptive text. While the generated text may be accurate and syntactically correct, it is often overly g…