1 paper
Arash Rocky, Q. M. Jonathan Wu
Vision-Language Models (VLMs) lag behind Large Language Models due to the scarcity of annotated datasets, as creating paired visual-textual annotations is labor-intensive and expen…