1 paper · 1 filter
Sri Harsha Dumpala, David Arps, Sageev Oore +2
Vision-language models (VLMs), serve as foundation models for multi-modal applications such as image captioning and text-to-image generation. Recent studies have highlighted limita…