2 papers
cs.CV2024
Seeing Syntax: Uncovering Syntactic Learning Limitations in Vision-Language Models
Sri Harsha Dumpala, David Arps, Sageev Oore +2
Vision-language models (VLMs), serve as foundation models for multi-modal applications such as image captioning and text-to-image generation. Recent studies have highlighted limita…
cs.CV2024
Sensitivity of Generative VLMs to Semantically and Lexically Altered Prompts
Sri Harsha Dumpala, Aman Jaiswal, Chandramouli Sastry +3
Despite the significant influx of prompt-tuning techniques for generative vision-language models (VLMs), it remains unclear how sensitive these models are to lexical and semantic a…