1 citations · 2 across the 9 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2024
CLoVe: Encoding Compositional Language in Contrastive Vision-Language Models
Santiago Castro, Amir Ziai, Avneesh Saluja +2
Recent years have witnessed a significant increase in the performance of Vision and Language tasks. Foundational Vision-Language Models (VLMs), such as CLIP, have been leveraged in…
cs.CV2023
Language Guided Visual Question Answering: Elevate Your Multimodal Language Model Using Knowledge-Enriched Prompts
Deepanway Ghosal, Navonil Majumder, Roy Ka-Wei Lee +2
Visual question answering (VQA) is the task of answering questions about an image. The task assumes an understanding of both the image and the question to provide a natural languag…