1 paper · 1 filter
Sergio Lanza, Jae Hee Lee, Stefan Wermter
Vision Language Models (VLMs) have demonstrated impressive performance in tasks requiring joint understanding of images and text, such as image captioning and Visual Question Answe…