2 citations · 2 across the 1 of their papers we have counts for
1 paper
Prateek Verma, Minh-Hao Van, Xintao Wu
Vision language models (VLMs) have recently emerged and gained the spotlight for their ability to comprehend the dual modality of image and textual data. VLMs such as LLaVA, ChatGP…